Claude Sonnet 5.5 vs Opus 5.5: When Half the Price Really Saves
Sonnet 5.5 costs half as much per token and suits well-scoped work at low or medium effort. Push it to max effort and it can cost more per task than Opus 5.5.
On this page

Run Claude Sonnet 5.5 at low or medium effort for well-scoped, checkable work: bug fixes, test-covered changes, documents, and high-volume tool loops. Keep Claude Opus 5.5 at its default medium effort for open-ended work that needs sustained judgment. If a task only goes well once you push Sonnet 5.5 to xhigh or max, move that task to Opus 5.5 instead. At that point Sonnet's half-price token rate no longer buys you a cheaper result.
Sonnet 5.5 launched on September 28, 2026, six days after Opus 5.5. Per million tokens it costs $2 input and $10 output, against $4 and $20 for Opus 5.5. Cache reads cost $0.20 on both (Anthropic pricing). That shared cache-read rate, together with how many tokens each model spends at a given effort level, decides whether "half the price" turns into a smaller bill. As of September 29, 2026, Anthropic's launch charts plot cost per task by effort level, and Artificial Analysis has measured both models at max effort only. Neither tells you what your own tasks will cost at medium, so the sections below show how to work that out.
Which model to start with, by kind of work
| Your work | Start with | Effort | Switch when |
|---|---|---|---|
| Clear spec, and tests or a checker can confirm the result (bug fixes, small features, refactors with tests) | Sonnet 5.5 | medium; low for simple edits | Failures keep recurring after one retry at high: send those tasks to Opus 5.5 |
| High-volume agent loops, PR review on every commit, bulk document or spreadsheet work | Sonnet 5.5 | low or medium | Misses start costing more than the tokens you saved |
| Chat and other latency-sensitive replies | Sonnet 5.5 | low or medium | Rarely; this is where Sonnet's speed matters most |
| Ambiguous specs, architecture calls, long multi-step work where a wrong early decision is expensive | Opus 5.5 | medium (its default on the API and in Claude Code) | Opus still misses at high or xhigh: see Claude Fable 5.1 vs Opus 5.5: Which Should You Use? |
| Hard code review where a missed bug is the costly failure | Opus 5.5 | medium or higher | Stays on Opus unless your own misses show otherwise |
| You don't know yet | Opus 5.5 | medium | Measure Sonnet 5.5 at low and medium against it, as shown at the end |
Most of these rows follow Anthropic's own positioning. Its launch post calls Sonnet 5.5 "strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets," while "Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment" (Introducing Claude Sonnet 5.5). The effort starting points come from Anthropic's Sonnet 5.5 prompting guide. For agentic coding it recommends medium for well-specified tasks and high for harder or longer ones. Its models overview still says "if you're unsure which model to use, start with Claude Opus 5.5 for most workloads" (models overview). That explains the last row.

A smaller option is also on the way. Anthropic says Claude Haiku 5.5, aimed at high-volume and cost-sensitive work, will join the family "in the coming weeks." No date or price has been announced.
Why "half the price" and "costs more per task" are both true
The price sheet on its own looks simple:
| Claude API, per million tokens | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache write, 5 minutes | $2.50 | $5 |
| Cache write, 1 hour | $4 | $8 |
| Cache read | $0.20 | $0.20 |
| Batch input / output | $1 / $5 | $2 / $10 |
| Fast mode (research preview) | Not offered | $8 / $40 |
Sonnet 5.5 costs exactly half on everything except cache reads, which are tied. Opus 5.5 gets a special cache-read rate of 0.05x its input price, while most models pay 0.1x. Sonnet 5.5 keeps Sonnet 5's price list unchanged, so the rates in Claude Sonnet 5 Price Increase Canceled: Current Rates and Billing Math still apply. Both models have a 1M-token context window and a 128K-token output limit, and a full 1M context is billed at the standard rates.
A per-task bill multiplies these prices by the tokens a model actually uses: thinking, output, extra turns, and the context it re-reads on each turn. Three things can erase Sonnet's price advantage.
Effort changes how much work Sonnet 5.5 does. Anthropic's launch post draws the line itself. Sonnet 5.5 "complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost." The prompting guide explains why the top settings get expensive. At xhigh and max, Sonnet 5.5 "can start its own rounds of review and verification, sometimes with subagents," and makes related fixes it noticed along the way. The benchmark table even carries a footnote for this: on FrontierCode, Sonnet 5.5 scored lower at max (46.2%) than at xhigh (52.1%). Anthropic says the reason was that at max it ran Claude Code's code-review skill more often, which led to timeouts and edits outside the task's scope.
Third-party measurements at max effort already show the reversal. Artificial Analysis runs both models at "Adaptive Reasoning, Max Effort" on its Intelligence Index v4.3.2 (10 evaluations). There, Sonnet 5.5 scored 56 at $7.60 per index task, and Opus 5.5 scored 58 at $5.98 per task. At max effort, the model with half the token price came out about 27% more expensive per task. Sonnet 5.5 produced 410M output tokens across the index, against 260M for Opus 5.5. Output tokens are only part of the gap, though. More turns also mean more input and more cache reads. These figures are from the Artificial Analysis Sonnet 5.5 page and Opus 5.5 page, a day after launch, and they cover max effort only.
The API defaults don't match. On the Claude API, Sonnet 5.5 defaults to high effort and Opus 5.5 defaults to medium. If you swap the model ID and leave effort unset, you are comparing Sonnet one level higher than Opus. In Claude Code, both models start at medium, so the mismatch only affects raw API calls. On the API, set effort explicitly on both sides before you compare anything.
The word "Max" can mean two things here. The Claude Max plan is a subscription tier. max effort is the highest effort level on the API and in Claude Code. The Artificial Analysis numbers above refer to the effort level.
Break-even: how much more work Sonnet 5.5 can do and still cost less
At identical token counts, Sonnet 5.5 costs half as much on everything except cache reads. That gives a simple way to see how much extra work it can do before the bill stops favoring it.
Split one task's Opus 5.5 bill into cache reads (R) and everything else (O). Run on Sonnet 5.5 with the same tokens, the task would cost R + O/2. Suppose Sonnet 5.5 needs k times as many tokens in every category, from more turns, more thinking, or more re-reading. Then it stays cheaper while:
k < (R + O) / (R + O/2) = 2 / (1 + y)
y = cache-read share of the task's Opus 5.5 bill| Cache reads as a share of your Opus 5.5 bill | Sonnet 5.5 stays cheaper until it uses... |
|---|---|
| 0% (no caching) | 2.0x the tokens |
| 25% | 1.6x |
| 50% | 1.33x |
| 75% | 1.14x |

Three illustrative token profiles, priced per turn at list rates (inputs are examples, not measured traffic):
Per turn: cache_read × $0.20 + fresh_input × input price + output × output price
(token counts in millions)
A. Uncached: 5k input, 6k output
Opus 5.5: 0.005 × 4 + 0.006 × 20 = $0.140
Sonnet 5.5: 0.005 × 2 + 0.006 × 10 = $0.070 → break-even 2.0x
B. Agent tool loop: 100k cache read, 5k fresh input, 6k output
Opus 5.5: 0.1 × 0.20 + 0.005 × 4 + 0.006 × 20 = $0.160
Sonnet 5.5: 0.1 × 0.20 + 0.005 × 2 + 0.006 × 10 = $0.090 → break-even ≈1.78x
C. Cache-heavy: 500k cache read, 2k fresh input, 3k output
Opus 5.5: 0.5 × 0.20 + 0.002 × 4 + 0.003 × 20 = $0.168
Sonnet 5.5: 0.5 × 0.20 + 0.002 × 2 + 0.003 × 10 = $0.134 → break-even ≈1.25xThe more of your bill that is cache reads, the less room Sonnet 5.5 has. Anthropic says cache reads "make up the majority of agentic and coding work costs" (in its Opus 5.5 announcement). If that holds for your agent, cache reads are above 50% of the Opus bill. Then Sonnet 5.5 only saves money while it needs less than about 1.33 times the tokens Opus 5.5 does. At low and medium effort that margin looks realistic, since Anthropic says Sonnet 5.5 costs less per task there. At max effort, the Artificial Analysis run above shows Sonnet going well past it.
The same ratio holds for batch jobs, because both models get the same 50% batch discount. Fast mode changes the comparison in the other direction. At $8 / $40 it doubles Opus 5.5's standard input and output rates, so an Opus fast-mode bill is being compared with Sonnet at a quarter of the price on those lines.
What the benchmarks settle, and what they don't
Anthropic's launch table puts the two models close on most rows. The table leaves out one thing that matters: most scores are at the highest effort each model was run at, not at the settings you would use day to day.
| Benchmark (Anthropic-reported) | Sonnet 5.5 | Opus 5.5 | Effort note |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% | Opus 5.5 at xhigh, its highest score |
| FrontierCode 1.1 (Main) | 52.1% at xhigh (46.2% at max) | 54.4% | Sonnet scored lower at max than at xhigh |
| CursorBench 4.0 | 55.5% | 57.8% | High-effort results |
| GDPval-AA v2.1 | 1844 | 1846 | Run by Artificial Analysis on a pre-release deployment |
| Humanity's Last Exam, with tools | 64.5% | 67.7% | High-effort results |
| OSWorld 2.1 (partial) | 80.1% | 81.8% | High-effort results |
The "beats Opus at coding" headline is the Terminal-Bench 4.0 row, where Sonnet 5.5 leads by about four points. On the other coding rows Opus 5.5 is ahead by two or three. At its default medium effort, Anthropic reports Opus 5.5 at 54.6% on FrontierCode and 52.5% on CursorBench. Don't set that 52.5% against Sonnet's 55.5% as if the two were run the same way. One is Opus at medium, the other is Sonnet at a high setting.
Independent coding results so far are small but consistent. CodeRabbit tested 13 hard pull-request cases, each with one verified bug. Sonnet 5.5 with thinking on caught 6. Opus 5.5 caught 8 at standard settings and 10 at max, both from CodeRabbit's earlier September run on the same cases. Opus also had higher actionable precision (66.7% at standard vs 41.2% for Sonnet 5.5). Sonnet 5.5's review calls cost about $0.46–$0.47 per review at list price, and CodeRabbit concluded that each model fits a different role. Opus 5.5 is for high-risk changes where a missed bug is expensive. Sonnet 5.5 is for the review pass on every pull request. CodeRabbit also ran one side-by-side Claude Code build. Sonnet 5.5 finished in 29 minutes 27 seconds and Opus 5.5 in 44 minutes 50 seconds, with results "close to identical" and Opus "a little higher in fidelity." CodeRabbit calls it "one run, not a benchmark," and 13 cases show a direction rather than proof.
In Claude Code and the Claude apps
On a Pro or Max subscription, you don't pay per token, so the practical trade-offs are speed and quality, plus how quickly you use up your plan's limits. Anthropic lists Sonnet 5.5's latency as "Fast" and Opus 5.5's as "Moderate." Sonnet 5.5 is Anthropic's fastest Sonnet so far, generating output more than 30% faster than Sonnet 5.
In Claude Code, Sonnet 5.5 and Opus 5.5 both start at medium effort. You can change the model and the effort level separately (Claude Code model configuration):
# Inside a session
/model # picker; left/right arrows set effort, Enter also saves the model as your default
/model sonnet # or /model opus; on the Anthropic API these aliases mean the 5.5 models
/effort high # low, medium, high, xhigh or max; /effort alone opens a slider, /effort auto clears it
/status # shows the current model
# One session only
claude --model claude-sonnet-5-5 --effort medium
claude --model claude-opus-5-5
# Every session started from this shell (zsh; use ~/.bashrc for bash)
echo 'export ANTHROPIC_MODEL="claude-sonnet-5-5"' >> ~/.zshrc
echo 'export CLAUDE_CODE_EFFORT_LEVEL="medium"' >> ~/.zshrcA few things to know before you rely on it:
- Plan access. Claude's pricing page lists Sonnet as available on Free, while Opus needs Pro ($20/month, or $17/month billed annually) or Max (from $100/month). The table doesn't name model versions, so check that Free shows Sonnet 5.5 in your model picker. For choosing between the paid plans, see Claude Pro vs Max in 2026: Pricing, Claude Code Limits, and When Max Is Worth It. Opus 5.5 also raised five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. Claude Opus 5.5 Pricing and Limit Reset: What Each One Changes covers that change.
- Changing effort is cheap; changing models is not. On these two models, with an API key or a Claude subscription, a new effort level keeps the prompt cache and applies right away. That makes
/effort highfor one hard step a low-cost move. It doesn't hold on Amazon Bedrock, Google Cloud, or a Claude apps gateway. Switching models is different. Each model has its own cache, so after/modelthe next request re-reads the whole conversation with no cache hits. Claude Code asks you to confirm while the cache is still warm (prompt caching in Claude Code). - Switching mid-session also drops the previous model's reasoning. On the API, Sonnet 5.5 can't read Opus 5.5's thinking blocks, and no other model reads Sonnet 5.5's. Blocks a model can't read are silently dropped. Claude Code's documentation doesn't spell out what
/modeldoes with earlier thinking, but the API rule suggests the new model continues without it. To hand a plan from Opus to Sonnet, have Opus write the plan into a file or a normal message first. Plain text carries over, and thinking blocks don't. opusplanis the built-in split. This alias usesopusin plan mode andsonnetfor execution. Each plan-mode toggle is a model switch, so it starts a fresh cache every time. Check what the aliases point to on your provider. On the Anthropic API,sonnetmeans Sonnet 5.5, but on Amazon Bedrock and Google Cloud it resolves to Sonnet 4.5, and on Claude Platform on AWS to Sonnet 4.6. On Microsoft Foundry,opusis Opus 4.6 andsonnetis Sonnet 4.5. On those providers, pin the full model ID, or setANTHROPIC_DEFAULT_SONNET_MODELto your provider's Sonnet 5.5 ID.- Safety fallbacks differ. In the Claude apps, Sonnet 5.5 is the first Sonnet with cyber fallbacks. A narrow set of flagged offensive-security requests and some frontier-AI-development requests are re-run on Sonnet 5, and the chat stays on Sonnet 5 until you switch back. Opus 5.5's cyber fallbacks go to Opus 4.8. Routine secure-coding work, such as scanning your own source code for vulnerabilities, stays on Sonnet 5.5 according to Anthropic's help page. The classifiers also read memory, connector content, search results, and files, so a switch can come from content you didn't type.
- Fast mode is Opus-only. Opus 5.5 fast mode is available in Claude Code and on the Claude Platform, and Anthropic claims "up to 2.5x speed." Sonnet 5.5 has no fast mode. It is fast by default.
Swapping model IDs in an API integration: what breaks
Both models reject request shapes that worked on Sonnet 5 and Opus 5. Moving between them has its own traps too. Check this list before you flip a model ID in production:
| Behavior | Sonnet 5.5 (claude-sonnet-5-5) | Opus 5.5 (claude-opus-5-5) |
|---|---|---|
| API default effort | high | medium |
thinking: {"type": "disabled"} | 400 error; send {"type": "between_tools"} to turn off up-front thinking | 400 error; thinking is always on, so lower effort instead |
between_tools limits | Only at low, medium, high; 400 at xhigh/max; no per-message effort changes | Not available |
Manual budget_tokens | 400 error | 400 error |
tool_choice any or tool | 400 error; use auto with strict: true or structured outputs | Same |
Non-default temperature, top_p, top_k | 400 error | 400 error (Opus 4.7 and later) |
| Text between tool calls | Arrives in thinking blocks, empty at default display: "omitted" | Same |
computer_20251124 tool | Rejected on Claude API and Google Cloud; use computer_toolset_20260801 (Bedrock still accepts it) | Same |
Fast mode (speed: "fast") | Not offered | Research preview, Claude API only |
Server-side fallback (fallbacks: "default", beta) | Retries cyber and frontier_llm refusals on Sonnet 5 | Retries on the model Anthropic recommends for the category |
| Thinking blocks after a switch | Can't read Opus 5.5 blocks; no other model reads Sonnet 5.5 blocks | Can't read Sonnet 5.5 blocks |
| Editing history before a thinking block | 400 by default on accounts created on or after August 31, 2026, unless you opt into drop_block | Same |
Sources: What's new in Claude Sonnet 5.5, What's new in Claude Opus 5.5, and the Opus 5.5 migration guide.
Two failures don't show up as errors. A streaming UI "goes quiet between tool calls, with no error" until you set thinking.display on either model. On Sonnet 5.5 you can also use between_tools to get the text back. And a router that sends one conversation to Opus 5.5 for some turns and to Sonnet 5.5 for others still gets successful requests. Each model simply works without the other's reasoning, and the dropped blocks aren't billed. If you route, route per task or per conversation, not per turn. On the API, keep top-level effort fixed within a cached conversation as well, since changing it invalidates the prompt cache. The per-message effort beta changes effort for a single turn while keeping the cache.
Using both: plan with Opus, build with Sonnet, escalate failures
"Opus plans, Sonnet implements" is a split some developers already use, and it can pay off. Only one version of it has measured support, though.
- Escalate failures. This has the strongest evidence. When a task can be checked automatically, run it cheap first and re-run only the failures on a stronger setting. Anthropic measured this within Opus 5.5 on a SWE-bench Pro subset. Running everything at
lowand re-running failures athighpassed about 97% of tasks for about $0.17 each. Running everything athighpassed 95.3% for $0.29 (Optimizing for cost and intelligence). Sending Sonnet 5.5's failures on to Opus 5.5 follows the same logic, but it hasn't been measured. It needs a reliable failure signal, and every failed first attempt adds latency. - Hand off a written plan. Opus 5.5 writes the plan into a file, and Sonnet 5.5 implements it in a new session. This avoids the dropped-thinking problem, because the plan travels as text.
opusplanin Claude Code. This is the same split, built in: Opus in plan mode, Sonnet for execution. It pays a full cache re-read at each plan-mode toggle, and Anthropic hasn't published cost data for it with the 5.5 models. Pin Sonnet 5.5 explicitly outside the Anthropic API, as described above.- Advisor tool (API beta). A Sonnet 5.5 executor can consult Opus 5.5 as its advisor, and the advice comes back encrypted. Opus 4.8, Opus 4.7, and Sonnet 5 advisors return a 400 with a Sonnet 5.5 executor. Anthropic's cost guide says advisors pay off when the advisor is priced "well above the executor" and actually gets consulted. Opus 5.5 is only twice Sonnet 5.5's price, and low-effort executors can stop asking. The guide's advice: "first price the advisor's model alone at low effort; that is the baseline to beat." Anthropic hasn't published a Sonnet 5.5 + Opus 5.5 advisor measurement.
Check it on your own traffic
Launch charts and max-effort indexes can't tell you what your tasks cost at the settings you actually run. An afternoon on your own tasks gives a more reliable answer:
- Pull 20–50 real tasks, weighted like production. Include the hardest tenth, because the cost of the tasks the cheaper model fails is what decides the bill. Anthropic's cost guide makes the same point about pricing the tail.
- Write an outcome check for each task: tests pass, ticket closed, expected rows returned.
- Run four arms in separate sessions: Sonnet 5.5 at
lowandmedium, and Opus 5.5 atlowandmedium. Set effort explicitly. Add Sonnet 5.5 athighifmediumfails too often. - Price every request from its
usageblock. Divide total spend, failed attempts included, by the number of tasks that passed.
PRICES = { # USD per million tokens, Claude API list prices as of September 29, 2026
"claude-sonnet-5-5": {"input": 2.00, "output": 10.00, "cache_read": 0.20},
"claude-opus-5-5": {"input": 4.00, "output": 20.00, "cache_read": 0.20},
}
def request_cost(model: str, usage: dict) -> float:
p = PRICES[model]
writes = usage.get("cache_creation") or {}
return (
usage["input_tokens"] * p["input"]
+ writes.get("ephemeral_5m_input_tokens", 0) * p["input"] * 1.25
+ writes.get("ephemeral_1h_input_tokens", 0) * p["input"] * 2.0
+ usage.get("cache_read_input_tokens", 0) * p["cache_read"]
+ usage["output_tokens"] * p["output"]
) / 1_000_000
def cost_per_accepted_task(model: str, tasks: list[dict]) -> float:
"""tasks: [{"usages": [usage, ...], "passed": bool}, ...]"""
spend = sum(request_cost(model, u) for t in tasks for u in t["usages"])
accepted = sum(t["passed"] for t in tasks)
return spend / accepted if accepted else float("inf")Read the result against the break-even table. If Sonnet 5.5 at medium passes about as many tasks as Opus 5.5 at medium, it wins as long as its token use stays under the break-even multiple for your cache-read share. If it only matches Opus at high or above, the savings are probably gone. In that case Opus 5.5 at medium is the simpler default. Anthropic's measurements are also "directional, not guarantees," and both models are days old, so re-run the check when prices or defaults change.
Frequently asked questions
Is Claude Sonnet 5.5 cheaper than Opus 5.5? Per token, yes: half the price on input, output, and cache writes, and the same $0.20 per million on cache reads. Per completed task, only when Sonnet 5.5 doesn't need much more work than Opus. At low and medium effort Anthropic says it costs less per task. At max effort, Artificial Analysis measured it at about 27% more per task than Opus 5.5.
Does Sonnet 5.5 really beat Opus 5.5 at coding?
On one benchmark. Anthropic reports Sonnet 5.5 at 70.6% on Terminal-Bench 4.0 against Opus 5.5's 66.4% at xhigh. Opus 5.5 leads on FrontierCode, CursorBench, and CodeRabbit's hard code-review cases. Anthropic itself says Opus stays stronger on open-ended work that needs sustained judgment.
Is Opus 5.5 cheaper than Opus 5?
Yes. Opus 5.5 is $4 / $20 per million input and output tokens, against $5 / $25 for Opus 5, and cache reads dropped 60% to $0.20. Anthropic says Opus 5.5 at default settings, where effort is now medium instead of Opus 5's high, costs 40% less than Opus 5 on typical workloads.
Which should I pick for my Claude Code default?
If most of your sessions are well-defined changes with tests, set Sonnet 5.5 as the default. Switch to Opus 5.5 with /model for design work and messy debugging, ideally at the start of a new task, because the switch re-reads the conversation without cache hits. opusplan automates the plan-then-build version. If most sessions are exploratory or span many files with unclear requirements, keep Opus 5.5 as the default. If you are also comparing vendors, Claude Opus 5.5 vs GPT-6 Sol: When the 2x Price Premium Pays Off covers that decision.





