Interleaved thinking lets Claude reason again after each tool call finishes, instead of only once before the first one — so a multi-step tool-calling loop can adjust its plan based on what a tool actually returned, rather than committing to a fixed sequence of calls up front. It’s controlled by a beta header or, on newer models, turned on automatically, and it changes both how you read the response and how prompt caching behaves across the loop.
What problem this actually solves
Without interleaved thinking, a model using extended thinking does its reasoning once, before it starts calling tools, and then executes on that plan. If a tool call midway through returns something unexpected — an empty result set, an error, a number that changes which branch makes sense — the model has no dedicated space to reason about it before deciding what to do next; it has to fold that judgment into the next tool call or the final answer without a visible reasoning step in between. Per Anthropic’s documentation, interleaved thinking changes that: “Claude can reason about the results of a tool call before deciding what to do next,” “chain multiple tool calls with reasoning steps in between,” and “make more nuanced decisions based on intermediate results.” The capability that changes isn’t whether tool calls can chain — regular tool use already supports calling several tools in sequence — it’s whether a visible reasoning step sits between each one.
Turning it on: the beta header, and which models need it
For models still using manual extended thinking (thinking: {type: "enabled"}), interleaved thinking requires sending the interleaved-thinking-2025-05-14 beta header. Support varies by model in a way that’s worth checking before you build around it: Claude Haiku 4.5 does not support it at all, and on the API “the beta header is accepted but ignored” rather than producing an error — which means a silent no-op is the actual failure mode to watch for, not a 400 response. Claude Opus 4.6 has a narrower limitation: its manual mode has no interleaved thinking at all, and only its adaptive mode interleaves, so getting reasoning between tool calls on that specific model means switching to adaptive thinking rather than adding the beta header.
That header-gated behavior is specific to manual-mode thinking. On adaptive thinking — the newer configuration style using thinking: {type: "adaptive"} instead of a fixed budget_tokens — interleaving is automatic wherever adaptive thinking itself is supported, with no header required. On every model in that adaptive-capable set, reasoning between tool calls simply appears in thinking blocks by default.
Adaptive versus manual: the migration you’re actually choosing
Anthropic’s docs describe the move from manual to adaptive thinking as a behavioral change, not just a different parameter shape: “With a fixed budget, Claude thinks on every request. With adaptive thinking, Claude determines whether and how much to think on each request, and at lower effort settings it may skip thinking entirely on easy inputs.” The practical migration looks like this — from a manual configuration setting a token budget directly:
{"model": "claude-sonnet-4-6", "max_tokens": 16000, "thinking": {"type": "enabled", "budget_tokens": 10000}}
to an adaptive configuration that hands the decision to the model, with an effort level instead of a token count:
{"model": "claude-sonnet-4-6", "max_tokens": 16000, "thinking": {"type": "adaptive"}, "output_config": {"effort": "high"}}
That trade-off cuts both ways depending on your workload. A fixed budget gives you a predictable upper bound on reasoning tokens and therefore cost per request, which matters for a high-volume endpoint with tight latency or spend targets. Letting the model decide whether to think at all is better suited to a mixed workload where most requests are simple and a fixed budget would be wasted reasoning on easy inputs, but it also means your cost per request becomes less predictable until you’ve measured it against your own traffic.
The budget_tokens rule that’s easy to get backwards
In ordinary manual-mode extended thinking, budget_tokens has a floor of 1,024 tokens and a ceiling set by max_tokens — since thinking tokens count toward that same per-turn limit, the budget has to leave room for the actual response. Interleaved thinking changes that second constraint: because the budget spans multiple thinking blocks across several tool calls within a single assistant turn rather than one block before a single response, budget_tokens is allowed to exceed max_tokens in that specific mode. Carrying the ordinary rule over into an interleaved-thinking setup — assuming the budget must stay under max_tokens the same way it does elsewhere — leads to a budget that’s artificially small for a loop that may run through several tool calls before finishing.
What you have to preserve when you pass tool results back
This is the detail most likely to break a working integration during a refactor. Anthropic’s documentation states it plainly: “when you return tool results, you must pass the thinking blocks from the assistant message back to the API, complete and unmodified.” With interleaved thinking, there are more of these blocks to track — one between each tool call rather than a single block at the start — and all of them have to round-trip back to the API unchanged, not just the final one before your last response. A common failure pattern is trimming or reformatting the conversation history before the next request, which silently drops or alters a thinking block and breaks a loop that otherwise looked correct in testing with short conversations, where there’s only one thinking block to get right.
How this interacts with prompt caching
Thinking blocks get cached along with tool results during a tool-use loop: once you send a follow-up request that includes tool results, the prior conversation history — thinking blocks included — becomes eligible for caching, and a cached thinking block still counts as input tokens in your usage metrics when it’s read back from cache. This happens automatically, without needing an explicit cache_control marker, and the mechanics are the same whether you’re using regular or interleaved thinking. The difference interleaved thinking introduces is in degree, not kind: because thinking blocks can now appear between several tool calls instead of just once, there are more of them in the conversation history, which means more opportunities for a change anywhere in that history to invalidate the cache for everything downstream of it. If your prompt caching strategy was tuned against a single-thinking-block conversation, it’s worth re-checking cache hit rates after turning on interleaved thinking rather than assuming the same cache breakpoints still make sense.
Watching what it actually costs
Because interleaved thinking can add a reasoning block between every tool call instead of just one before the first call, the token cost of a single turn is easy to underestimate from reading the request payload alone — the budget you set describes a ceiling across the whole turn, not what any one step will use. The response carries the actual figure: Anthropic’s extended-thinking documentation points to the usage.output_tokens_details.thinking_tokens field, which “reports how many of the billed output tokens were internal reasoning,” as the place to check rather than estimating from the visible thinking text. Logging that field per request during testing — before the loop goes anywhere near production traffic — is a cheap way to catch a tool-calling sequence that’s spending far more on reasoning than the task justifies, which is easy to miss when you’re only looking at whether the final answer came back correct. If the tool outputs feeding that reasoning are themselves structured data, tightening up how you request structured JSON output from the tool calls can also cut down on reasoning spent parsing a messy result before the model can act on it.
When to reach for this versus when not to
Interleaved thinking earns its added complexity on agentic workflows where a later tool call genuinely depends on judgment about an earlier one — a research agent that decides which source to check next based on what the last search turned up, or a debugging agent that picks its next diagnostic step based on what the last command’s output showed. It adds little for a single-tool-call request, or for a fixed sequence of tool calls that doesn’t actually change based on intermediate results — in both of those cases, the extra reasoning blocks are overhead without a corresponding gain in decision quality, and a simpler manual-mode setup without the beta header is easier to reason about and debug.
A migration checklist if you’re adding this to an existing tool-use loop
- Confirm your model actually supports it — check for Haiku 4.5 (unsupported, header silently ignored) and Opus 4.6 manual mode (unsupported; switch to adaptive) before assuming the header alone is enough.
- If you’re on manual-mode thinking, add the
interleaved-thinking-2025-05-14beta header rather than just raisingbudget_tokens— the header is what actually changes where thinking blocks appear, not the budget size. - Re-check your
budget_tokensceiling: it’s allowed to exceedmax_tokenshere, unlike ordinary extended thinking, so a value copied over from a non-interleaved setup may be unnecessarily conservative. - Audit every place your code trims, summarizes, or reformats conversation history before a follow-up request — each thinking block has to be passed back complete and unmodified, and a loop with several tool calls now has several of these to preserve instead of one.
- Re-measure your prompt cache hit rate after enabling it, rather than assuming your existing cache breakpoints still behave the same way with more thinking blocks in the history.
FAQ
Does interleaved thinking let the model chain tool calls that it couldn’t chain before?
No — chaining itself isn’t the capability being added. Per Anthropic’s documentation, “Claude can chain tool calls with or without interleaved thinking. Interleaving changes where thinking blocks appear between tool calls, not whether tool calls can chain.” What changes is whether there’s a visible reasoning step between each call, not whether the calls themselves can be sequenced.
If I move to adaptive thinking, do I still need the interleaved-thinking beta header?
No. On every model that supports adaptive thinking, interleaved thinking is automatic once adaptive thinking is enabled, with no beta header involved — the header only matters for models still using the manual, fixed-budget configuration.
Photo: Sillerkiil / Wikimedia Commons (CC BY-SA 4.0)



