💡 TL;DR / Summary — Claude Opus 5.5 Pricing (BLUF)
- Question: Is Claude Opus 5.5 actually cheaper than Opus 5, now that its thinking can’t be switched off?
- Measurement: From Anthropic’s official price list, the same token counts cost 20% to 35.5% less on Opus 5.5 across three workload shapes. The saving is larger when more of the input is cached.
- Catch: Thinking is billed as output. If Opus 5.5 produces more than +41.7% output for an uncached request (+32.3% for a cached session), it costs more than Opus 5.
- Rule: Switch cache-heavy agent work first; measure
thinking_tokensbefore moving workloads that ran Opus 5 with thinking off.
Anthropic released Claude Opus 5.5 on September 22, 2026 at $4 per million input tokens and $20 per million output tokens, 20% below Opus 5, with cache reads cut 60% to $0.20. The announcement also states that Opus 5.5 “is also no longer available with ’thinking’ mode switched off.”
Anthropic’s thinking documentation counts reasoning inside the billed output tokens, at the output rate. A model with lower per-token prices can therefore still cost more per task if it generates more reasoning. This post uses the official price list to calculate the cost of three workload shapes on each model, and the output increase at which Opus 5.5 stops being cheaper.
What exactly changed in the price list?
| Item (per 1M tokens) | Claude Opus 5 | Claude Opus 5.5 | Change |
|---|---|---|---|
| Base input | $5 | $4 | −20% |
| Output | $25 | $20 | −20% |
| Cache read (hit) | $0.50 | $0.20 | −60% |
| 5-minute cache write | $6.25 | $5 | −20% |
| 1-hour cache write | $10 | $8 | −20% |
| Batch input / output | $2.50 / $12.50 | $2 / $10 | −20% |
| Fast mode input / output | $10 / $50 | $8 / $40 | −20% |
| Thinking | On by default; can be disabled at effort high or below | Always on; cannot be disabled | — |
The cache-read cut is three times deeper than the input cut. Anthropic prices Opus 5.5 cache hits at 0.05x base input instead of the standard 0.1x. Thinking also changed: the July 24 release notes for Opus 5 allowed disabling thinking at effort high or below, while the September 22 notes for Opus 5.5 describe “always-on adaptive thinking” that cannot be turned off.
Both models share the newer tokenizer introduced with Claude 4.7, so the same text produces the same token counts on either model; the comparison below does not need a tokenizer adjustment.
How much does the same work cost on each model?
We priced three workload shapes with identical token counts on both models. The first two reuse the token mix from Anthropic’s own worked example on the pricing page (50,000 input tokens and 15,000 output tokens per session, with 40,000 of the input cached in the second case). The third is a long agent loop where most of the context is re-read from cache every turn.
| Workload shape | Opus 5 | Opus 5.5 | Saving | Sonnet 5 (reference) |
|---|---|---|---|---|
| Plain request: 50k input uncached, 15k output | $0.625 | $0.500 | −20.0% | $0.250 |
| Cached coding session: 10k uncached + 40k cache reads, 15k output | $0.445 | $0.348 | −21.8% | $0.178 |
| Long agent loop: 50k uncached + 950k cache reads, 20k output | $1.225 | $0.790 | −35.5% | $0.490 |
The arithmetic, so you can rerun it with your own token counts:
PRICE = { # $ per 1M tokens: uncached input, cache read, output
"opus-5": (5.00, 0.50, 25.00),
"opus-5.5": (4.00, 0.20, 20.00),
"sonnet-5": (2.00, 0.20, 10.00),
}
def cost(model, uncached, cached, output):
i, c, o = PRICE[model]
return (uncached * i + cached * c + output * o) / 1_000_000
for shape in [(50_000, 0, 15_000), (10_000, 40_000, 15_000), (50_000, 950_000, 20_000)]:
print(shape, [round(cost(m, *shape), 3) for m in PRICE])
# (50000, 0, 15000) [0.625, 0.5, 0.25]
# (10000, 40000, 15000) [0.445, 0.348, 0.178]
# (50000, 950000, 20000) [1.225, 0.79, 0.49]
The saving grows with the share of input served from cache, because cache reads carry Opus 5.5’s largest discount. Anthropic’s announcement claims Opus 5.5 “costs 40% less than Opus 5 on typical workloads.” List-price arithmetic alone lands between 20% and 35.5% on these shapes; any gap up to 40% would have to come from Opus 5.5 using fewer tokens for the same task, which this post does not measure.
When does always-on thinking erase the saving?
Because thinking tokens are billed as output, the question becomes: how much more output can Opus 5.5 generate for the same task before it costs more than Opus 5? Holding input fixed and solving for the output multiplier gives one break-even point per workload shape.
| Workload shape | Opus 5 cost | Opus 5.5 fixed input cost | Break-even output (vs Opus 5) |
|---|---|---|---|
| Plain request | $0.625 | $0.200 | 1.42x (+41.7%) |
| Cached coding session | $0.445 | $0.048 | 1.32x (+32.3%) |
| Long agent loop | $1.225 | $0.390 | 2.09x (+109%) |
Example: if a task on Opus 5 ran with thinking disabled and produced 15,000 output tokens, Opus 5.5 can spend up to about 6,250 extra tokens on reasoning (15,000 × 0.417) before a plain request costs more than it did on Opus 5. In a cached session the cushion is thinner, about 4,850 tokens (15,000 × 0.323), because the input side has already been discounted almost to nothing. In a cache-heavy agent loop, output can roughly double before the saving disappears.
Two further effects are not measured here. Adaptive thinking decides per request whether to think, and at lower effort settings Anthropic’s documentation says Claude “may skip thinking entirely on easy inputs,” which lowers the cost of simple calls. Claude models from Opus 4.5 onward keep prior turns’ thinking blocks in context and bill them as input, so in long multi-turn sessions earlier reasoning is billed again as input on later turns.
Which workloads should move to Opus 5.5 first?
| Workload | Move to Opus 5.5? | Why |
|---|---|---|
| Cache-heavy agent loops (long sessions re-reading a codebase) | Yes, first | Largest list-price saving (−35.5%) and the widest cushion before thinking erases it (+109%) |
| Opus 5 work that already ran with thinking on | Yes | Same behavior, 20% lower rates on every token class |
Opus 5 work that ran with thinking off at effort high or below | Measure first | Forced thinking adds output; the saving holds only below +32% to +42% extra output |
| Latency-sensitive paths on fast mode | Yes, if you stay on the Claude API | Fast mode drops from $10/$50 to $8/$40; not available on partner clouds or with Batch |
| High-volume everyday coding | Consider Sonnet 5 instead | $2/$10 is half of Opus 5.5 on input and output |
A practical checklist before switching a production path:
- Log
usage.output_tokens_details.thinking_tokenson a sample of real requests to both models. That field shows how many output tokens went to reasoning. - Compare output per task, not per request. If Opus 5.5 finishes in fewer turns, a higher per-turn output can still win.
- Hold the effort level constant across a cached conversation. Changing thinking configuration mid-conversation invalidates cache breakpoints, which removes the 60% cache-read discount.
- Re-check the price page before committing. Anthropic has changed Claude prices twice since June: Sonnet 5’s introductory $2/$10 became permanent, and Opus 5.5 undercut Opus 5.
FAQ
Q. How much does Claude Opus 5.5 cost?
$4 per million input tokens and $20 per million output tokens on the Claude API. Cache reads are $0.20 per million, 5-minute cache writes $5 and 1-hour cache writes $8. The Batch API halves input and output to $2 and $10.
Q. Is Opus 5.5 cheaper than Opus 5?
On list price, yes: input and output are 20% lower and cache reads 60% lower. For identical token counts, our three workload shapes cost 20% to 35.5% less, with the biggest saving on cache-heavy agent loops.
Q. Can I turn off thinking on Claude Opus 5.5?
No. The release notes describe always-on adaptive thinking that cannot be disabled, and the announcement says Opus 5.5 is no longer available with thinking switched off. Opus 5 allowed disabling thinking at effort high or below.
Q. Are thinking tokens billed as output?
Yes. Anthropic’s thinking documentation counts reasoning inside the billed output tokens and breaks it out in usage.output_tokens_details.thinking_tokens. On Opus 5.5 that reasoning costs the $20 per million output rate.
Q. When does Opus 5.5 end up costing more than Opus 5?
When its output for the same task grows past the break-even point: about +41.7% for a plain uncached request, +32.3% for a cached coding session, and +109% for a long cache-heavy agent loop.
Q. How does Opus 5.5 compare with Sonnet 5 on price?
Sonnet 5 lists at $2/$10 with $0.20 cache reads, half of Opus 5.5 on input and output. In the three workload shapes above, Sonnet 5 costs $0.25, $0.178 and $0.49 against Opus 5.5’s $0.50, $0.348 and $0.79.
Update log
- 2026-09-28 — Initial publication. All prices verified the same day against Anthropic’s official pricing page, the Opus 5.5 announcement and the Claude Platform release notes. Workload costs are arithmetic on those list prices, not billing measurements.
Pricing in this category changes frequently. Figures reflect Anthropic’s official sources as of September 28, 2026; re-verify against the pricing page before budgeting.
Related Research
- Claude Fable 5 shutdown and restoration: the June–July export-control episode, and the current Claude model lineup with list prices.
Sources: Anthropic — Claude API Pricing · Anthropic — Introducing Claude Opus 5.5 · Anthropic — Claude Platform Release Notes · Anthropic — Extended Thinking
Last verified: September 2026
Key figures
Frequently asked
How much does Claude Opus 5.5 cost?
Is Opus 5.5 cheaper than Opus 5?
Can I turn off thinking on Claude Opus 5.5?
Are thinking tokens billed as output?
When does Opus 5.5 end up costing more than Opus 5?
How does Opus 5.5 compare with Sonnet 5 on price?
Does fast mode cost more on Opus 5.5?
Sources
- Anthropic — Claude API pricing (official)platform.claude.com
- Anthropic — Introducing Claude Opus 5.5www.anthropic.com
- Anthropic — Claude Platform release notesplatform.claude.com
- Anthropic — Extended thinking (billing of thinking tokens)platform.claude.com
Educational content only — not investment or financial advice. Data, prices, and tool specifications change; verify independently and paper-trade before risking capital.
