If Claude is sitting on "almost done thinking" and you're wondering whether something broke — it didn't. That status means Claude is in extended thinking (its reasoning mode): it's planning the answer before it starts writing. A few seconds, even 20–30 seconds, of this is normal. What is not normal is waiting minutes with nothing coming back. Those are two different problems, and this guide separates them and gives you the fix for each.
What "almost done thinking" actually means
"Almost done thinking" is the label Claude shows while it's running its extended reasoning pass. Instead of replying token-by-token immediately, the model spends a budget of "thinking" tokens working out a plan, then produces the visible answer. This is the same mechanism behind "thinking with high effort" in Claude Code and the reasoning indicators in the Claude apps.
The clearest way to read it: the phrase is a progress signal, not an error. As one r/ClaudeCode thread puts it, the extended-reasoning pause is "Claude planning before executing, not a server issue." So when you see claude almost done thinking, the model is working — the only question is whether it's working too long.
A rough line to draw:
Seconds to ~30s of thinking → normal, especially for hard reasoning or coding tasks.
Minutes with no output, repeatedly → something is wrong; skip to the fixes below.
Why it takes so long (or hangs entirely)
Slowness and a true hang have different causes. Three things drive the slow case:
Extended-thinking depth. On a quick lookup, Claude can lean toward more reasoning effort than the task needs — it "thinks" hard about a question that didn't need it.
Sequential tool calls. In agentic use (Claude Code), most of the wall-clock time isn't the model reasoning — it's tool calls. One breakdown of Claude Code latency measured each file read, search, or test run at roughly 300–800ms as a synchronous round-trip, and notes they don't run in parallel by default — so a vague prompt that triggers a dozen-plus exploratory calls stacks those round-trips into ten-plus seconds before any real work happens.
Context bloat. The entire transcript is re-sent to the model every turn. As a session fills up, responses slow down and quality drops — the same analysis observed a noticeable slowdown once a session passes roughly 60% of the context window, though the exact point varies (it's the "lost in the middle" effect on details buried deep in the conversation).
The true hang is a separate failure. A tracked Claude Code issue (#32526) describes new sessions getting stuck on "thinking" and never producing output — no error, even to a plain "hello" — while a previously-open session keeps working fine. That report came from a heavy setup: many PreToolUse hooks, multiple MCP servers, 80+ registered skills, and a custom (Bedrock) provider. If yours never returns a single token, treat it as a hang, not slowness.
How to fix it
Quick fixes (try these first)
/clearto drop the conversation and start clean — the fastest cure for context bloat./compactto summarize and shrink the context. Note it's lossy by design, so save anything important to a file first.Restart the session, or hop back to an older session that's still responding — the workaround most users land on for a hang.
Control the effort level
The most overlooked speed lever is effort. Claude tends to default to high/max reasoning; matching effort to the task gives large speed gains with no quality loss on routine work. A working cheat sheet:
Task | Effort | Why |
|---|---|---|
Quick lookup, summary, formatting |
| No deep reasoning needed; near-instant |
Standard coding, drafting |
| Balanced |
Architecture, hard debugging, math |
| Worth the wait |
If the opacity bothers you (Claude hides thinking detail by default), Claude Code can surface a thinking summary so you can at least see what it's doing.
When it's truly stuck, not just slow
If you get zero output, it's a hang, not depth:
Trim the startup load — temporarily disable extra
PreToolUsehooks, unused MCP servers, and skills, then reopen the session.Check your provider — custom model IDs and gateways (Bedrock and similar) show up in a number of hang reports.
Run with
--verboseto see what it's actually doing: a long string of file reads points to a tool-call problem; a slow first response with no tool calls points to latency or context.
Advanced: control thinking from the API
The apps give you limited control over thinking. The API gives you that control directly — and that's the practical fix if you need predictable latency. The same effort levels from the cheat sheet above are an API parameter, and you can also switch extended thinking off entirely:
message = client.messages.create(
model="claude-opus-4-6",
max_tokens=4096,
thinking={"type": "adaptive"}, # Claude decides how much to think
output_config={"effort": "low"}, # low | medium | high | max — caps the depth
# or, to skip extended thinking entirely:
# thinking={"type": "disabled"},
messages=[{"role": "user", "content": "..."}],
)On current Claude models (Opus 4.6 and up) you don't set a fixed token budget — you set an effort level (low for quick work, up to max), or disable extended thinking outright. That's the same lever the apps hide, exposed as a parameter you control. (See the official extended-thinking docs for the current reference — parameter names can shift between SDK versions, so check the one you're on.)
Any Anthropic-compatible endpoint can issue these calls — the official API, or a compatible mirror like AIReiter, where the request above works unchanged. Which one you use matters less than the takeaway: thinking control lives at the API layer, and the apps don't expose it.
Is Claude "getting worse"?
This is the question lurking behind most "why is claude almost done thinking forever" searches, and the honest answer is: usually it's not permanent. Much of the perceived regression traces to default behavior changes — server-side adjustments to thinking budgets, or conservative defaults rolled out in updates — rather than the model getting dumber. There's a lot of community back-and-forth on this, and the recurring conclusion is the same: the experience recovers once you take back control — set the effort level, clear bloated sessions, and give structured prompts. If Claude feels worse this week, change those three things before concluding it's broken.
FAQ
What does "almost done thinking" mean?
It means Claude is in its extended-reasoning mode, working out a plan before it writes the answer — a normal progress state, not an error. It only signals a problem when it never resolves.
Why does Claude take so long to think?
Three usual causes: a high default effort level, slow sequential tool calls (~300–800ms each) in agentic sessions, and a bloated context window. Lowering the effort level, reducing the number of tool calls, and clearing context each help.
Does Claude ever get stuck thinking?
Yes — distinct from slowness. New sessions can hang on "thinking" and never return output, often tied to heavy hook/MCP/skill setups or custom providers. Restart the session or return to a working one.
Claude isn't finishing the response — what should I do?
Treat it as a hang: /clear or restart, trim startup load, and check your provider. If it's slow rather than stalled, lower the effort level and shrink the context.
Can I make Claude think faster?
Yes. Set effort to low/medium for routine tasks, keep sessions short, and for full control call the API with a lower effort level or extended thinking disabled.