Skip to main content

Claude Sonnet 5.5 vs GPT-6 Astra: When the Cheaper Model Wins

Sonnet 5.5 beats GPT-6 Astra only at max effort, at a higher cost per task. Astra leads from low to high, ties at xhigh; Sonnet's low rates win on long inputs.

LaoZhang AI TeamPublished16 min read
On this page
Claude Sonnet 5.5 vs GPT-6 Astra cover comparing list prices of $2 and $10 against $10 and $50 per million tokens, with max-effort scores of 56 and 53 and costs per task of $7.60 and $3.26

Claude Sonnet 5.5 is neither a straight upgrade over GPT-6 Astra nor a cheap imitation of it. Its list rates are one-fifth of Astra's. On Artificial Analysis' Intelligence Index, though, it beats Astra only at max effort, and at that setting it spends so many tokens that each task costs about 2.3 times what Astra's does. At low, medium, and high effort, Astra scores higher. Sonnet 5.5's price advantage is still real, and it widens on requests above 272,000 input tokens, where OpenAI reprices the entire Astra request.

In practice: keep Astra for terminal, SRE, and scientific agent work, and wherever finishing a hard task quickly at top effort matters. Trial Sonnet 5.5 for long-document and very large-context requests, for finance, tax, and legal-research tasks where it led Vals AI's max-effort runs, and for short interactive turns where its first token arrives sooner. If what you actually want is OpenAI quality at Sonnet's price, the like-for-like model is GPT-6 Sol, not Astra.

Prices below are USD per million tokens on each vendor's standard API tier, as of September 29, 2026. Benchmark numbers are third-party measurements, each named with its tester and settings. Sonnet 5.5 launched on September 28, 2026, so expect these figures to move.

Which model to start with for each kind of work

WorkloadStart withWhyWhat would change the answer
Terminal, SRE, and scientific-workflow agentsGPT-6 AstraVals AI (both at max effort): Astra ahead by 27.14 points on Terminal-Bench Science and 26.72 points on SRE BenchYour own trial; testers disagree on Terminal-Bench 4.0
Finance, tax, legal research, medical codingTrial Sonnet 5.5Vals AI (max): Sonnet ahead on Tax Agent Bench (+10.06), Legal Research Bench (+8.65), Finance Agent (+4.56), MedCode (+4.43)At max effort Sonnet cost more per test on some of these and took roughly 2–4x as long
Requests above 272K input tokensSonnet 5.5No long-context surcharge; Astra bills the whole request at $20 input and $75 outputWhether Sonnet's answers pass your acceptance bar on those documents
Short interactive turns and chatSonnet 5.5 at low or mediumArtificial Analysis: 0.92 s vs 2.42 s to first answer token at low; cheapest cost per taskIts index scores there (36 and 41) sit below Astra's lowest setting (46)
Hardest problems where the score matters more than the billSonnet 5.5 at max, or Claude Opus 5.5Artificial Analysis: Sonnet max 56 vs Astra max 53; Opus 5.5 max scores 58About 2.3x Astra's cost per task, about 7x the output tokens, longer runs
Cheapest strong model while staying on OpenAIGPT-6 SolSame list price as Sonnet 5.5, no API migrationYour trial of Sol against Sonnet 5.5

What the "Sonnet 5.5 beats Astra" headlines actually measure

Both models take a reasoning effort setting with five levels: low, medium, high, xhigh, and max. Effort controls how much the model thinks before and between its answers. Higher effort usually buys a better score at the cost of more output tokens and more time. Each vendor defines its own levels, so Sonnet's high is not the same amount of work as Astra's high. Comparing "same label" pairs is a convenient reading, not a calibrated one.

Artificial Analysis ran both models at every level. Its numbers, captured on September 29, 2026:

EffortSonnet 5.5 scoreSonnet 5.5 cost per taskGPT-6 Astra scoreGPT-6 Astra cost per taskAstra ÷ Sonnet cost
low36$0.4146$0.822.0x
medium41$0.5950$1.542.6x
high47$1.0851$1.731.6x
xhigh52$2.7452$2.310.84x
max56$7.6053$3.260.43x

Scores are the Artificial Analysis Intelligence Index v4.3.2 (a 10-evaluation suite). Cost per task is what one index task cost to run at list prices. The last column is simple division of the two cost columns.

Bar chart of Artificial Analysis Intelligence Index scores and cost per task for Claude Sonnet 5.5 and GPT-6 Astra at each effort level from low to max

Four things stand out.

  1. The headline is max against max. Sonnet 5.5 at max scores 56, above Astra's best of 53, and ranks second overall behind Claude Opus 5.5 at max (58). To get there it used about 193,000 output tokens per index task, roughly 7x what Astra used at max. A 5x cheaper token cannot absorb a 7x larger token count, so Sonnet's task cost ends up at $7.60 against Astra's $3.26.
  2. From low to high, Astra leads; at xhigh they tie. At matching labels Astra is ahead by 10 points at low, 9 at medium, and 4 at high. The two tie at xhigh (52), where Astra is also cheaper per task.
  3. At similar scores, Astra is often cheaper. Astra at low scores 46 for $0.82 per task. Sonnet needs high effort to reach 47, at $1.08. Artificial Analysis draws the same conclusion: Sonnet 5.5 "sits off the Intelligence vs. Cost per Task Pareto Frontier," and at lower efforts "GPT-6 Astra or Sol configurations" deliver "equivalent performance for lower cost."
  4. Where Sonnet is cheaper, it scores lower. Sonnet at low and medium is the cheapest option in the table per task, but its 36 and 41 are below every Astra setting.

Defaults matter here because most people never change them. Sonnet 5.5 defaults to high on the Claude API (47 on this index) and to medium in the Claude apps and Claude Code (41). OpenAI's documentation doesn't state a default effort for Astra, so set it explicitly in any comparison.

One caveat on the Sonnet numbers: Artificial Analysis tested a pre-release deployment that had a bug affecting structured-output requests. Anthropic expects "minimal change or slightly understated performance" after the fix, and Artificial Analysis plans to re-run the affected evaluations. About 0.1% of tasks fell back to Sonnet 5 under Anthropic's default fallback setting.

Vals AI at max effort

Vals AI ran both models at max effort across 19 shared benchmarks, as of September 29, 2026. Sonnet 5.5 scored higher on 13 and Astra on 6. The reported ±1 standard-error ranges overlap on 9 of the 19, and Vals states that this "is not a pairwise statistical significance test," so several of those 13 wins are close calls.

  • Astra's clear leads: Terminal-Bench Science (65.71% vs 38.57%), SRE Bench (56.87% vs 30.15%), and IOI competitive programming (100% vs 83.06%). Its category averages are higher for coding (78.60% vs 74.58%) and science (66.04% vs 56.26%).
  • Sonnet's clear leads: Tax Agent Bench (73.39% vs 63.34%), Legal Research Bench (48.08% vs 39.42%), and Finance Agent (58.10% vs 53.54%). It averages higher in finance, healthcare, legal, and education.
  • Results that shift without fallback: Sonnet's CyberBench score (59.58% vs Astra's 41.07%) leans on Sonnet 5 answering tasks that Sonnet 5.5 declined. Fallback handled 54 of 116 tasks, and counting those as failures drops Sonnet to 41.97%, level with Astra. The same adjustment takes Sonnet's SRE Bench score from 30.15% to 19.08%, widening Astra's lead there.
  • Overall: Vals Index 69.22% for Sonnet vs 66.61% for Astra. Sonnet ran with server-side fallback to Sonnet 5; counting fallback tasks as failures lowers its Vals Index to 68.89%, still ahead.

The cost and time columns tell the same story as Artificial Analysis. Running the Vals Index cost $20.80 on Sonnet 5.5 and $19.09 on Astra, and took 1 hour 10 minutes versus 25 minutes 11 seconds. At max effort, Sonnet's lower rates bought no savings on the full suite, and every benchmark with a recorded time finished sooner on Astra.

Don't pair numbers from the two launch pages

Anthropic's Sonnet 5.5 page compares it with Sonnet 5, Opus 5.5, and GPT-6 Sol; Astra isn't in the table. OpenAI's Astra page has no Sonnet 5.5 column. Headlines that set Anthropic's Terminal-Bench 4.0 figure for Sonnet 5.5 (70.6%) against OpenAI's figure for Astra at max (57.9%) are mixing two vendors' test setups. Independent runs of that benchmark disagree even on direction: Artificial Analysis measured Sonnet max at 64% and Astra xhigh at 60%, while Vals measured Sonnet at 53.03% and Astra at 57.07%. Treat Terminal-Bench 4.0 as unsettled between these two models. The vendor OSWorld numbers aren't comparable either, because they come from different versions of the benchmark (2.1 and 2.0).

Anthropic's own speed and cost claims are relative to Sonnet 5, not Astra: Sonnet 5.5 "runs 30%+ faster, and costs up to 30% less for most work." Artificial Analysis measured Sonnet 5.5 at max costing about 50% more per index task than Sonnet 5 at max, so the "up to" depends heavily on effort and workload.

Price: 5x per token, rarely 5x per task

Rate, USD per 1M tokensClaude Sonnet 5.5GPT-6 Astra, up to 272K inputGPT-6 Astra, above 272K input
Input$2$10$20
Cached input (cache read)$0.20$1$2
Cache write$2.50 (5-minute), $4 (1-hour)$12.50$25
Output$10$50$75
Batch$1 input / $5 output50% of standard50% of standard

Anthropic applies no long-context surcharge anywhere in Sonnet 5.5's 1M-token window: "A 900k-token request is billed at the same per-token rate as a 9k-token request." OpenAI bills any Astra prompt above 272K input tokens at "2x input and cache rates and 1.5x output for the full request," not just for the tokens past the threshold. Astra's Flex tier is also half price, and its Fast mode (formerly Priority) doubles the standard rate. Sonnet 5.5 has no fast mode. Region-pinned processing adds 10% on both sides: US-only inference for Sonnet 5.5 is billed at 1.1x, and OpenAI charges a 10% uplift for data-residency endpoints. For the full Astra rate card, see GPT-6 Astra API Pricing: Rates, Caching, and Cost Examples. Sonnet 5.5 bills exactly like Sonnet 5, so Claude Sonnet 5 Price Increase Canceled: Current Rates and Billing Math covers the cache reconciliation details.

Up to 272K input, every line is exactly 5x. Prompt caching therefore doesn't narrow the gap; it lowers both bills in the same proportion. Only two things move the ratio at list prices: crossing 272K input, which pushes Astra to 10x on input and 7.5x on output, and the number of tokens each model actually spends on the job.

Worked examples on identical token counts

Formula: cost = (fresh input × input rate + cached input × cache-read rate + cache writes × cache-write rate + output × output rate) ÷ 1,000,000. Output includes reasoning tokens. The examples assume the same token counts on both models, standard tier, and no tool fees.

RequestClaude Sonnet 5.5GPT-6 AstraAstra ÷ Sonnet
A. 100K fresh input + 20K output$0.20 + $0.20 = $0.40$1.00 + $1.00 = $2.005x
B. Agent turn: 250K cached + 15K fresh input, 8K output (265K input)$0.05 + $0.03 + $0.08 = $0.16$0.25 + $0.15 + $0.40 = $0.805x
C. Same turn after context grows to 290K cached (305K input)$0.058 + $0.03 + $0.08 = $0.168$0.58 + $0.30 + $0.60 = $1.48≈8.8x
D. 300K fresh input + 20K output$0.60 + $0.20 = $0.80$6.00 + $1.50 = $7.50≈9.4x

Rows B and C show why long agent sessions deserve a look. The request grows by only 40,000 cached tokens, yet Astra's cost nearly doubles because the whole request moves to the higher tier. Sonnet's cost barely changes.

Four worked price examples on identical token counts showing GPT-6 Astra's cost ratio jumping from 5x to about 9x once a request passes 272K input tokens, while Claude Sonnet 5.5 stays at a flat rate

These are arithmetic on list prices, not bills. The two models use different tokenizers, so the same text yields different token counts. They also spend different amounts on reasoning and retries. On Artificial Analysis' tasks, that difference in spend turned a 5x rate gap into anything from Astra costing 2.6x more per task (medium) to Sonnet costing 2.3x more (max). The only cost number that settles the question is your own cost per accepted task, covered in the trial section below.

If the real question is "cheapest good model," compare GPT-6 Sol

GPT-6 Sol costs $2 input, $0.20 cached input, $2.50 cache write, and $10 output up to 272K input, the same as Sonnet 5.5. That is why Anthropic's launch table benchmarks Sonnet 5.5 against Sol rather than Astra. Artificial Analysis rates Sonnet 5.5 at high effort as its most competitive setting on cost, "very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task." Sol does have a long-context tier ($4 / $0.40 / $5 / $15 above 272K input), which Sonnet 5.5 doesn't.

If your stack is on OpenAI and the appeal of Sonnet 5.5 is the price, test Sol first: it needs no API migration. GPT-6 Sol vs Terra vs Luna vs Astra: Which Model Should You Use? covers when to step down from Astra within OpenAI's lineup. Going the other way, if you want the top score and can pay for it, Claude Opus 5.5 at max leads the Artificial Analysis index at 58. Claude Opus 5.5 vs GPT-6 Sol: When the 2x Price Premium Pays Off works through that pairing.

ChatGPT and Claude subscriptions

API prices say nothing about subscription allowances, so treat the two separately.

  • ChatGPT: "GPT-6 Pro, powered by GPT-6 Astra" is available in chat on the $100 and $200 Pro plans, Business, and Enterprise. Plus includes Astra in ChatGPT Work and Codex, but not GPT-6 Pro in chat. Free and Go don't include Astra. Codex usage for Astra is published as estimated local messages per five hours (for Plus, 5–45), which OpenAI describes as "not fixed message limits."
  • Claude apps: Sonnet 5.5 is available in Claude on web, desktop, and mobile, in Cowork, and in Claude Code, with effort set to medium by default. Anthropic's support and pricing pages don't list which paid plans include it. When Sonnet 5.5 flags a conversation as cyber-security or AI-model-development content, the app can switch it to Sonnet 5; biology and distillation-type requests are blocked.

On the app side, the default matters more than the model name. Sonnet 5.5 at medium scored 41 on Artificial Analysis' index, below Astra's lowest setting. If you're comparing Claude Code against Codex on hard tasks, raise Sonnet's effort before drawing conclusions.

What changes when you move API traffic from Astra to Sonnet 5.5

Moving between vendors is a rewrite of the request layer, not a model-ID swap. The list below combines OpenAI's and Anthropic's migration documentation.

  1. Different API. Astra runs on OpenAI's Responses and Chat Completions APIs; its tool calling requires Responses. Sonnet 5.5 uses Anthropic's Messages API with its own tool schema. Both models are on Amazon Bedrock, which can simplify accounts and billing but doesn't change the request shape.
  2. Forced tool use returns 400. tool_choice set to any or a named tool is rejected on Sonnet 5.5. Keep auto, set strict: true for schema-valid inputs, and say in the prompt when the tool applies.
  3. No "thinking off" switch. thinking: {"type": "disabled"} and manual budget_tokens return 400. The lowest setting is thinking: {"type": "between_tools"}, which works only at low, medium, or high effort.
  4. Sampling parameters and prefill. Any non-default temperature, top_p, or top_k returns 400 on Sonnet 5.5, and so does a prefilled final assistant turn. Astra already requires you to drop temperature and top_p because it has no none effort level, so prompts written for Astra usually carry over on this point.
  5. Streaming goes quiet between tool calls. Sonnet 5.5 returns its notes between tool calls as thinking blocks, which are empty under the default display: "omitted". A UI that streams those notes shows nothing, with no error, until you change display or use between_tools.
  6. Keep conversation history append-only. Sonnet 5.5 thinking blocks are tied to the model, the conversation, and the account. For accounts created on or after August 31, 2026, replaying a thinking block after editing the system prompt, tools, or an earlier message returns 400 by default. Change instructions with mid-conversation system messages instead.
  7. Caching works differently. Anthropic charges explicit cache writes with 5-minute or 1-hour lifetimes and caches prompts of 512 tokens or more. Astra has its own cache-write price and a "30m" cache TTL option. Re-measure your cache hit rate rather than assuming it carries over.
  8. Re-run the effort sweep. Effort levels don't map across vendors, and Anthropic says Sonnet 5.5's levels were recalibrated even relative to Sonnet 5. Anthropic suggests starting at high in general, and at medium for well-specified agentic coding tasks.
  9. Tools that exist on only one side. Astra's Responses API tool list includes hosted_shell, code_interpreter, apply_patch, and image_generation. If your pipeline depends on one of them, plan a Claude-side equivalent or run it yourself. On Sonnet 5.5, computer use on the Claude API and Google Cloud requires the computer_toolset_20260801 toolset.
  10. Refusals and fallback. Sonnet 5.5 can decline in five stop_details categories and returns HTTP 200 with stop_reason: "refusal". Server-side fallback (beta) retries cyber and frontier_llm declines on Sonnet 5, which matters mainly for security-research workloads.

If Claude access itself is the obstacle, How to Get Stable Claude Access: Anthropic Direct, Supported Cloud, or the laozhang.ai Gateway? compares the routes.

Run a cost-per-accepted-task trial before switching

Cost per accepted task is the total billed cost of a run divided by the number of tasks whose output met your acceptance rule. It folds in token prices, token volume, retries, and failures, which is what the vendor tables and per-token comparisons leave out.

  1. Pick 20–50 real tasks from your workload and write down the acceptance rule for each: tests pass, a reviewer signs off, or the extracted fields match.
  2. Run both models on identical inputs and tools at a stated effort pair. Start with Sonnet 5.5 high against Astra medium or high, then move each one level up. Artificial Analysis places Sonnet high (47) close to Astra low (46), so that pair is worth a run too.
  3. Record billed usage per task: fresh input, cached input, cache writes, and output including reasoning tokens. Also log retries, wall-clock time, and minutes of human repair.
  4. Compute cost per accepted task for each model and effort level.
  5. Switch a workload only if Sonnet 5.5 clears the same acceptance bar at a lower cost per accepted task, with a turnaround you can live with.

This script turns usage counts into cost, including Astra's 272K rule. Replace the rates if your tier differs (halve both for Batch, or halve Astra for Flex).

python
# USD per 1M tokens, standard tier, as of September 29, 2026
RATES = {
    "claude-sonnet-5-5": {"input": 2.00, "cached": 0.20, "cache_write": 2.50, "output": 10.00},
    "gpt-6-astra":       {"input": 10.00, "cached": 1.00, "cache_write": 12.50, "output": 50.00},
    "gpt-6-astra-long":  {"input": 20.00, "cached": 2.00, "cache_write": 25.00, "output": 75.00},
}

def request_cost(model, fresh, cached=0, cache_write=0, output=0):
    """Output must include reasoning tokens. Uses the 5-minute cache-write rate for Sonnet."""
    key = model
    if model == "gpt-6-astra" and fresh + cached + cache_write > 272_000:
        key = "gpt-6-astra-long"  # whole request repriced above 272K input
    r = RATES[key]
    return (fresh * r["input"] + cached * r["cached"]
            + cache_write * r["cache_write"] + output * r["output"]) / 1_000_000

def cost_per_accepted_task(runs):
    """runs: list of dicts with model, fresh, cached, cache_write, output, accepted."""
    total = sum(request_cost(r["model"], r["fresh"], r["cached"],
                             r["cache_write"], r["output"]) for r in runs)
    accepted = sum(1 for r in runs if r["accepted"])
    return total / accepted if accepted else float("inf")

print(request_cost("claude-sonnet-5-5", 300_000, output=20_000))  # 0.8
print(request_cost("gpt-6-astra", 300_000, output=20_000))        # 7.5

Include every retry in runs, not just the final attempt, so failed attempts count toward the total cost.

Common questions

Which GPT model is the equivalent of Claude Sonnet 5.5?

By price, GPT-6 Sol: both list at $2 input and $10 output per million tokens up to 272K input, and Anthropic's own launch comparison uses Sol. GPT-6 Astra is OpenAI's top model at five times Sonnet's rates, so a Sonnet 5.5 vs Astra comparison is a cheaper model against a flagship.

Is Claude Sonnet 5.5 better than GPT-6 Astra for coding?

It depends on the kind of coding and the effort level. Vals AI's max-effort runs put Astra ahead on its coding category average, IOI, and SRE Bench, and Sonnet ahead on Vibe Code Bench and Code Migration by margins inside the reported uncertainty. Terminal-Bench 4.0 results point in opposite directions depending on the tester. Run your own tasks before moving an agent pipeline.

Will these benchmark numbers change?

Almost certainly. Artificial Analysis tested a pre-release Sonnet 5.5 deployment and plans to re-run affected evaluations, and both Artificial Analysis and Vals AI update their pages as models and harnesses change. The prices are first-party list rates as of September 29, 2026.

Workload decision board comparing GPT-6 Astra and Claude Fable 5.1 across access, API cost, data rules, and acceptance testing
Model Comparisons

GPT-6 Astra vs Claude Fable 5.1: Choose by Workload

Astra and Fable share the same headline input and output rates, but cache pricing, long-context billing, access, and safeguards can reverse the cheaper or safer choice.

7 min