💡 TL;DR / Summary — Claude Opus 5.5 Pricing (BLUF)

  • Question: Is Claude Opus 5.5 actually cheaper than Opus 5, now that its thinking can’t be switched off?
  • Measurement: From Anthropic’s official price list, the same token counts cost 20% to 35.5% less on Opus 5.5 across three workload shapes. The saving is larger when more of the input is cached.
  • Catch: Thinking is billed as output. If Opus 5.5 produces more than +41.7% output for an uncached request (+32.3% for a cached session), it costs more than Opus 5.
  • Rule: Switch cache-heavy agent work first; measure thinking_tokens before moving workloads that ran Opus 5 with thinking off.

Anthropic released Claude Opus 5.5 on September 22, 2026 at $4 per million input tokens and $20 per million output tokens, 20% below Opus 5, with cache reads cut 60% to $0.20. The announcement also states that Opus 5.5 “is also no longer available with ’thinking’ mode switched off.”

Anthropic’s thinking documentation counts reasoning inside the billed output tokens, at the output rate. A model with lower per-token prices can therefore still cost more per task if it generates more reasoning. This post uses the official price list to calculate the cost of three workload shapes on each model, and the output increase at which Opus 5.5 stops being cheaper.

What exactly changed in the price list?

Item (per 1M tokens)Claude Opus 5Claude Opus 5.5Change
Base input$5$4−20%
Output$25$20−20%
Cache read (hit)$0.50$0.20−60%
5-minute cache write$6.25$5−20%
1-hour cache write$10$8−20%
Batch input / output$2.50 / $12.50$2 / $10−20%
Fast mode input / output$10 / $50$8 / $40−20%
ThinkingOn by default; can be disabled at effort high or belowAlways on; cannot be disabled—

The cache-read cut is three times deeper than the input cut. Anthropic prices Opus 5.5 cache hits at 0.05x base input instead of the standard 0.1x. Thinking also changed: the July 24 release notes for Opus 5 allowed disabling thinking at effort high or below, while the September 22 notes for Opus 5.5 describe “always-on adaptive thinking” that cannot be turned off.

Both models share the newer tokenizer introduced with Claude 4.7, so the same text produces the same token counts on either model; the comparison below does not need a tokenizer adjustment.

How much does the same work cost on each model?

We priced three workload shapes with identical token counts on both models. The first two reuse the token mix from Anthropic’s own worked example on the pricing page (50,000 input tokens and 15,000 output tokens per session, with 40,000 of the input cached in the second case). The third is a long agent loop where most of the context is re-read from cache every turn.

Workload shapeOpus 5Opus 5.5SavingSonnet 5 (reference)
Plain request: 50k input uncached, 15k output$0.625$0.500−20.0%$0.250
Cached coding session: 10k uncached + 40k cache reads, 15k output$0.445$0.348−21.8%$0.178
Long agent loop: 50k uncached + 950k cache reads, 20k output$1.225$0.790−35.5%$0.490

The arithmetic, so you can rerun it with your own token counts:

PRICE = {  # $ per 1M tokens: uncached input, cache read, output
    "opus-5":   (5.00, 0.50, 25.00),
    "opus-5.5": (4.00, 0.20, 20.00),
    "sonnet-5": (2.00, 0.20, 10.00),
}

def cost(model, uncached, cached, output):
    i, c, o = PRICE[model]
    return (uncached * i + cached * c + output * o) / 1_000_000

for shape in [(50_000, 0, 15_000), (10_000, 40_000, 15_000), (50_000, 950_000, 20_000)]:
    print(shape, [round(cost(m, *shape), 3) for m in PRICE])
# (50000, 0, 15000)       [0.625, 0.5, 0.25]
# (10000, 40000, 15000)   [0.445, 0.348, 0.178]
# (50000, 950000, 20000)  [1.225, 0.79, 0.49]

The saving grows with the share of input served from cache, because cache reads carry Opus 5.5’s largest discount. Anthropic’s announcement claims Opus 5.5 “costs 40% less than Opus 5 on typical workloads.” List-price arithmetic alone lands between 20% and 35.5% on these shapes; any gap up to 40% would have to come from Opus 5.5 using fewer tokens for the same task, which this post does not measure.

When does always-on thinking erase the saving?

Because thinking tokens are billed as output, the question becomes: how much more output can Opus 5.5 generate for the same task before it costs more than Opus 5? Holding input fixed and solving for the output multiplier gives one break-even point per workload shape.

Workload shapeOpus 5 costOpus 5.5 fixed input costBreak-even output (vs Opus 5)
Plain request$0.625$0.2001.42x (+41.7%)
Cached coding session$0.445$0.0481.32x (+32.3%)
Long agent loop$1.225$0.3902.09x (+109%)

Example: if a task on Opus 5 ran with thinking disabled and produced 15,000 output tokens, Opus 5.5 can spend up to about 6,250 extra tokens on reasoning (15,000 × 0.417) before a plain request costs more than it did on Opus 5. In a cached session the cushion is thinner, about 4,850 tokens (15,000 × 0.323), because the input side has already been discounted almost to nothing. In a cache-heavy agent loop, output can roughly double before the saving disappears.

Two further effects are not measured here. Adaptive thinking decides per request whether to think, and at lower effort settings Anthropic’s documentation says Claude “may skip thinking entirely on easy inputs,” which lowers the cost of simple calls. Claude models from Opus 4.5 onward keep prior turns’ thinking blocks in context and bill them as input, so in long multi-turn sessions earlier reasoning is billed again as input on later turns.

Which workloads should move to Opus 5.5 first?

WorkloadMove to Opus 5.5?Why
Cache-heavy agent loops (long sessions re-reading a codebase)Yes, firstLargest list-price saving (−35.5%) and the widest cushion before thinking erases it (+109%)
Opus 5 work that already ran with thinking onYesSame behavior, 20% lower rates on every token class
Opus 5 work that ran with thinking off at effort high or belowMeasure firstForced thinking adds output; the saving holds only below +32% to +42% extra output
Latency-sensitive paths on fast modeYes, if you stay on the Claude APIFast mode drops from $10/$50 to $8/$40; not available on partner clouds or with Batch
High-volume everyday codingConsider Sonnet 5 instead$2/$10 is half of Opus 5.5 on input and output

A practical checklist before switching a production path:

  1. Log usage.output_tokens_details.thinking_tokens on a sample of real requests to both models. That field shows how many output tokens went to reasoning.
  2. Compare output per task, not per request. If Opus 5.5 finishes in fewer turns, a higher per-turn output can still win.
  3. Hold the effort level constant across a cached conversation. Changing thinking configuration mid-conversation invalidates cache breakpoints, which removes the 60% cache-read discount.
  4. Re-check the price page before committing. Anthropic has changed Claude prices twice since June: Sonnet 5’s introductory $2/$10 became permanent, and Opus 5.5 undercut Opus 5.

FAQ

Q. How much does Claude Opus 5.5 cost?

$4 per million input tokens and $20 per million output tokens on the Claude API. Cache reads are $0.20 per million, 5-minute cache writes $5 and 1-hour cache writes $8. The Batch API halves input and output to $2 and $10.

Q. Is Opus 5.5 cheaper than Opus 5?

On list price, yes: input and output are 20% lower and cache reads 60% lower. For identical token counts, our three workload shapes cost 20% to 35.5% less, with the biggest saving on cache-heavy agent loops.

Q. Can I turn off thinking on Claude Opus 5.5?

No. The release notes describe always-on adaptive thinking that cannot be disabled, and the announcement says Opus 5.5 is no longer available with thinking switched off. Opus 5 allowed disabling thinking at effort high or below.

Q. Are thinking tokens billed as output?

Yes. Anthropic’s thinking documentation counts reasoning inside the billed output tokens and breaks it out in usage.output_tokens_details.thinking_tokens. On Opus 5.5 that reasoning costs the $20 per million output rate.

Q. When does Opus 5.5 end up costing more than Opus 5?

When its output for the same task grows past the break-even point: about +41.7% for a plain uncached request, +32.3% for a cached coding session, and +109% for a long cache-heavy agent loop.

Q. How does Opus 5.5 compare with Sonnet 5 on price?

Sonnet 5 lists at $2/$10 with $0.20 cache reads, half of Opus 5.5 on input and output. In the three workload shapes above, Sonnet 5 costs $0.25, $0.178 and $0.49 against Opus 5.5’s $0.50, $0.348 and $0.79.

Update log

  • 2026-09-28 — Initial publication. All prices verified the same day against Anthropic’s official pricing page, the Opus 5.5 announcement and the Claude Platform release notes. Workload costs are arithmetic on those list prices, not billing measurements.

Pricing in this category changes frequently. Figures reflect Anthropic’s official sources as of September 28, 2026; re-verify against the pricing page before budgeting.


Sources: Anthropic — Claude API Pricing · Anthropic — Introducing Claude Opus 5.5 · Anthropic — Claude Platform Release Notes · Anthropic — Extended Thinking

Last verified: September 2026

Key figures

$4 / $20 per 1M tokens
Opus 5.5 list price
$0.20 per 1M (60% below Opus 5)
Cache reads
+32.3% to +109% depending on cache share
Break-even output increase

Frequently asked

How much does Claude Opus 5.5 cost?
$4 per million input tokens and $20 per million output tokens on the Claude API, with cache reads at $0.20 per million, 5-minute cache writes at $5 and 1-hour cache writes at $8. Batch processing halves input and output to $2 and $10.
Is Opus 5.5 cheaper than Opus 5?
On list price, yes. Input and output are 20% lower ($4/$20 against $5/$25) and cache reads are 60% lower ($0.20 against $0.50). Across our three workload shapes, the same token counts cost 20% to 35.5% less, with the largest saving in cache-heavy agent loops.
Can I turn off thinking on Claude Opus 5.5?
No. Anthropic’s release notes describe always-on adaptive thinking that cannot be disabled, and the announcement says Opus 5.5 is no longer available with thinking switched off. Opus 5 allowed disabling thinking at effort high or below.
Are thinking tokens billed as output?
Yes. Anthropic’s thinking documentation reports reasoning inside the billed output tokens, broken out in the usage.output_tokens_details.thinking_tokens field. Extra thinking therefore costs the $20 per million output rate on Opus 5.5.
When does Opus 5.5 end up costing more than Opus 5?
When its output grows past a break-even point for the same task. For a plain uncached request that point is about 1.42x Opus 5’s output (+41.7%); for a cached coding session about 1.32x (+32.3%); for a long cache-heavy agent loop about 2.09x (+109%).
How does Opus 5.5 compare with Sonnet 5 on price?
Sonnet 5 lists at $2/$10 with $0.20 cache reads, so on input and output it is half the price of Opus 5.5. In our three workload shapes Sonnet 5 costs $0.25, $0.178 and $0.49 against Opus 5.5’s $0.50, $0.348 and $0.79.
Does fast mode cost more on Opus 5.5?
Fast mode on Opus 5.5 lists at $8 input and $40 output per million tokens, against $10 and $50 for Opus 5 fast mode. It is available on the Claude API only, not on partner cloud platforms, and not with the Batch API.

Sources

  1. Anthropic — Claude API pricing (official)platform.claude.com
  2. Anthropic — Introducing Claude Opus 5.5www.anthropic.com
  3. Anthropic — Claude Platform release notesplatform.claude.com
  4. Anthropic — Extended thinking (billing of thinking tokens)platform.claude.com

Educational content only — not investment or financial advice. Data, prices, and tool specifications change; verify independently and paper-trade before risking capital.