With OpenAI's August 26, 2026 shutdown at the door, the highest-risk mistake is treating the migration as a rename. The Assistants API sunset was announced a full year ahead, in OpenAI's deprecation notice of August 26, 2025, and the replacement is the Responses API. Object names map cleanly on paper; the orchestration underneath them does not, and developers who followed the official guide still shipped breakage. Below: what dies, what the mappings hide, and what to do with whatever runway you have left.
August 26, 2026: what stops working, and what survives
Every Assistants endpoint family returns errors after the deadline. The shutdown covers /v1/assistants, /v1/threads, thread messages, runs, and run steps, including anything still sending the OpenAI-Beta: assistants=v2 header. Assistant configurations and thread history become unreachable through the API.
Not everything attached to an Assistants integration dies with it:
| Gone on August 26, 2026 | Still available |
|---|---|
/v1/assistants CRUD endpoints | Vector stores and uploaded files, reusable via Responses file search |
/v1/threads, thread messages | Chat Completions API (not part of this shutdown) |
| Runs and run steps | Responses API and Conversations API |
OpenAI-Beta: assistants=v2 workflows | Realtime API |
OpenAI's own deprecation tracker lists Responses plus Conversations as the designated replacements:
The four object mappings — and the two footnotes that change your architecture
OpenAI's migration guide maps four Assistants concepts onto Responses-era equivalents:
| Assistants API | Replacement | What actually changes |
|---|---|---|
Assistants | Prompts | Config moves to a dashboard-created, versioned object |
Threads | Conversations | Stores generalized items (messages, tool calls, tool outputs), not messages alone |
Runs | Responses | The create-run-poll-retrieve loop collapses into one responses.create call |
Run steps | Items | A union type covering messages, function calls, and results |
The loop collapse shows up in the official examples: a completed run on gpt-4.1 reports 34 prompt tokens and 130 completion tokens, while a completed response on gpt-5.5 reports 17 input tokens and 150 output tokens — same shape of workload, different field names.
That rename is footnote number one. Billing dashboards and payload parsers keyed to the old usage fields fail silently when the names change:
| Assistants field | Responses field |
|---|---|
usage.prompt_tokens | usage.input_tokens |
usage.completion_tokens | usage.output_tokens |
max_completion_tokens / max_prompt_tokens | max_output_tokens |
truncation_strategy | truncation |
object: "thread.run" | object: "response" |
Footnote number two is architectural. Prompts can only be created in the dashboard, not through an API, which breaks any system that dynamically creates one Assistant per customer, workspace, or document set. The official guide itself advises checking the prompt deprecation timeline before adopting prompt objects in a long-lived integration, because reusable prompt objects carry their own sunset risk. The durable pattern is instructions, tool schemas, and model choice held in your own source control and passed with each request. On thread history, OpenAI's position is one sentence: "We will not provide an automated tool for migrating Threads to Conversations."
Migrating the three built-in tools
Each Assistants tool has a concrete landing spot in Responses, and each shifts work into your application:
| Assistants tool | Where it lands in Responses | What your application now owns |
|---|---|---|
| File search | Vector stores survive; supply vector_store_ids in the tool definition at request time | Resolving the right store IDs before each call |
| Code interpreter | Container configured with type: "auto" | Container lifecycle |
| Functions | Nested function key removed; name, description, parameters move up one level | The tool loop: execute the call, return the result with its matching call_id, decide whether to loop |
The file-search row is the quiet architectural change for multi-tenant apps. One vector store per tenant used to be a setup-time binding on the Assistant object; now whichever tenant owns the incoming session must resolve to the right store IDs before the request goes out.
Failure reports from teams that already migrated
OpenAI's stated rationale is that Responses reached feature parity. The migration reports below describe parity at the object level, with a real refactor underneath. A multi-tenant chatbot SaaS owner documented a two-week migration on r/aiagents, and the failures survived a faithful reading of the official guide:
I had to retrofit every optional field as
["type", "null"]which feels like a type system workaround. — u/aidenclarke_12
Strict tool schemas require optional properties to be declared as nullable and still listed in required, so schemas grow and each handler that assumed missing-means-absent needs a second look. The same developer flagged where the deeper shift lives:
the vector store plumbing change is the real architectural shift. — u/aidenclarke_12
Streaming is the second silent breaker. Assistants run streaming does not adapt to Responses; it gets rewritten against typed server-sent events like response.created, response.output_text.delta, response.completed, and response.function_call_arguments.delta / .done, with explicit completion events and new tool-call event shapes — the event names are catalogued in migration coverage. SSE proxies and client handlers both need the rewrite, including reconnect logic.
The third breaker is ecosystem lag rather than the API itself:
the response API is out for a long time, but still a lot of frameworks SDKs do not support that — u/zhlmmc
If your stack sits on an agent framework that still assumes the Threads/Runs model (the lag u/zhlmmc ran into), budget time for that layer as well as your own glue code.
Choosing a state strategy: chain, Conversations, or manual replay
Three ways to keep multi-turn context in Responses, and they are not interchangeable:
| Strategy | Best for | Watch out for |
|---|---|---|
previous_response_id | Simplest chaining, minimal rewrites | Prior context stays in billable input |
| Conversations API | Closest analogue to Threads; server-side history | Backfill is yours to build; no vendor tool |
Manual replay, store: false | ZDR and strict retention requirements | You own all state; reasoning items must be carried forward |
For history, OpenAI's recommended sequence for converting an old thread:
- List the thread's messages in ascending order.
- Convert each user text message to
input_text. - Convert each assistant text message to
output_text. - Convert image URL content to
input_image, preservingimage_urlanddetail. - Create the Conversation with the converted
items.
Getting a role mapping wrong has a specific failure mode: the model reads its own past answers as fresh user instructions. Stored responses carry a 30-day default TTL unless you pass store: false, and conversations sit outside that response TTL with no separately published duration as of late July 2026, per the migration coverage that tracked it — a detail that matters if your disclosures promise a deletion window.
What the migration does to your token bill
Two billing facts matter.
First, previous_response_id is convenience, not a discount: OpenAI's Responses migration guide states that previous input tokens in the response chain are still billed as input tokens, so long-running conversations grow linearly unless you prune.
Second, cached input runs far cheaper than uncached: roughly a tenth of the input rate across the GPT-5.x tiers as listed in July 2026, and 40–80% better cache utilization on Responses versus Chat Completions in OpenAI-reported internal testing, per the compiled coverage. Treat the utilization range as a vendor number until your own dashboards confirm it. The parity check that matters is your own token count per session, before and after cutover.
If the migration is your moment to reprice the GPT-5.x workload itself, the GPT-5.6 pricing breakdown covers the per-token math, and OpenAI-compatible endpoints such as the GPT-5.6 API page run the same Responses-style workloads for direct comparison.
A migration plan sized to the runway you have left
1–6 days left. Back up first: list assistants and vector stores with limit=100, pull your files, serialize SDK objects with model_dump(). Backup-first write-ups flag the hard limit: there is no list-threads endpoint, so you can only export thread IDs your own application already stored. Then cut over behind a flag: new sessions go to Responses immediately, old threads get lazily backfilled only when a user reopens them.
A week or more. Convert one low-risk flow end to end before touching the rest. Rebuild the tool loop and verify each function result carries its matching call_id; replace stream handling with event-type branching; then compare behavior, latency, token usage, and error rates against the Assistants baseline before expanding traffic.
Past the deadline. The endpoints error and assistant configurations are gone from the API side; recovery means rebuilding from whatever your application database and backups hold, with vector stores and files still reachable through file search.
The unresolved trade-off: you swap a server-managed lifecycle (polling, truncation, the tool loop) for a single-call model with orchestration you can see and test. One developer who shipped on both put the exchange this way:
The responses API is the perfect meet in the middle—it manages heavy lifting but it's still flexible enough to manage for own functionality. — u/landongarrison
OpenAI Assistants API shutdown FAQ
Is the Chat Completions API also shutting down?
No. Chat Completions is not part of the August 26, 2026 shutdown, and OpenAI's guidance treats it as migratable to Responses one flow at a time rather than on a forced deadline.
Will OpenAI automatically migrate my existing threads?
No. The official migration guide states plainly: "We will not provide an automated tool for migrating Threads to Conversations." Backfill is application code you write, following the item-conversion sequence above.
Can I keep using the Assistants API after August 26, 2026?
No. Assistants, threads, messages, runs, and run steps all return errors after the date, including assistants=v2 workflows. Export anything you need before the deadline.
Do stored responses expire?
Yes. Stored responses default to a 30-day retention window unless you pass store: false, with conversations sitting outside that TTL as of the July 2026 reporting.
Do I have to move my assistant configuration into Prompts?
No, and for dynamically generated assistants you should not. Prompts are dashboard-created only, and the official guide itself flags reusable prompt objects for deprecation review; holding instructions and tool schemas in source control and passing them per request is the durable pattern.