Gemini 4 Argon vs Claude Opus 5.5 vs GPT-6.1 Sol: What to Run Now
Gemini 4 Argon is not open to developers yet. At matched high effort it ties Opus 5.5 and costs 6.2x GPT-6.1 Sol per task, despite the same $2/$10 list rate.
On this page

Of the three, only Claude Opus 5.5 and GPT-6.1 Sol can go into a project today. Google announced Gemini 4 Argon on September 30, 2026, but as of October 1 it is rolling out only to security teams in Google's Fairwind Program. The announcement gives no date, no model ID and no rate limits for developers. "GPT 6.1" in this matchup means GPT-6.1 Sol, the model OpenAI released on September 29.
The headline price needs one correction before you plan around it. Argon's introductory rate of $2 input and $10 output per million tokens matches GPT-6.1 Sol and is half of Opus 5.5. That holds only for identical token counts. On Artificial Analysis' index at matched "high" effort, Argon cost $1.99 per task, Opus 5.5 cost $1.82 and GPT-6.1 Sol cost $0.32, with scores of 53, 54 and 50. Argon ties Opus on both score and cost there, and it costs about six times what Sol does for three more points.
The benchmark figures below come from Google, Artificial Analysis and Vals AI, each under its own settings, and those settings change the answer more than the model names do.
Which of the three you can call today
| Gemini 4 Argon | Claude Opus 5.5 | GPT-6.1 Sol | |
|---|---|---|---|
| Status as of October 1, 2026 | Announced September 30; limited rollout | Released September 22; generally available | Released September 29; generally available |
| Who can use it | Trusted cyber defenders in the Fairwind Program | Claude Pro, Max, Team and Enterprise; Claude Code; Claude API; AWS, Google Cloud and Azure | OpenAI API; ChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu |
| Model ID | Not published | claude-opus-5-5 | gpt-6.1-sol |
| Context window | 1M per Artificial Analysis and Vals; not stated by Google | 1M | 1,050,000 |
| Maximum output | 1M announced; 262K per Vals | 128K | 128,000 |
Google's announcement says the company is taking part in the U.S. government's voluntary process for pre-release model access "while we gradually expand access." It promises availability to developers, enterprises and consumers "as soon as possible," starting with paid API customers and Google AI Ultra subscribers. No date is attached to any of those steps. The Gemini API pricing page listed Gemini 3.8 Flash and Gemini 3.1 Pro Preview on October 1, with no Argon or Gemini 4 entry.
Fairwind is Google's program for vetted security defenders, and Google says those teams get Argon without cyber guardrails. If you work on a security team, the Fairwind eligibility guide explains what the program is and who qualifies. That guide was written for Gemini 3.8 Flash Cyber, and Google has not published a separate application path for Argon.
Two limits apply on the OpenAI side. GPT-6.1 Sol is not in ordinary ChatGPT chat or on the Free and Go plans, and tool calls require the Responses API. The GPT-6.1 Sol pricing and access guide covers what breaks when you switch from the older Sol.
Same token price as Sol, half of Opus, and why that misleads
List rates for identical tokens
All rates are USD per million tokens on each vendor's standard tier.
| Argon, introductory | Argon, after the introductory period | GPT-6.1 Sol, up to 272K input | GPT-6.1 Sol, above 272K input | Opus 5.5 | |
|---|---|---|---|---|---|
| Input | $2 | $4 | $2 | $4 | $4 |
| Cached input | $0.10 (95% off) | Not published | $0.10 | $0.20 | $0.20 |
| Output | $10 | $20 | $10 | $15 | $20 |
| Cache write | Not published | Not published | $2.50 | $5 | $5 (5-minute) |
| Batch discount | Not published | Not published | 50% | 50% | 50% |
With U as fresh input, R as cache reads and O as output including reasoning, all in millions of tokens, the bill for one request is:
- Argon, introductory: 2U + 0.10R + 10O
- GPT-6.1 Sol up to 272K input: 2U + 0.10R + 10O, plus 2.5 per million tokens written to cache
- Opus 5.5: 4U + 0.20R + 20O, plus 5 per million tokens written to the 5-minute cache
- Argon after the introductory period: 4U + 20O, plus a cache rate Google has not published
| Request, before cache-write fees | Argon, introductory | GPT-6.1 Sol | Opus 5.5 | Argon, later rate |
|---|---|---|---|---|
| 100K fresh input, 20K output | $0.40 | $0.40 | $0.80 | $0.80 |
| 50K fresh input, 150K cache reads, 20K output | $0.315 | $0.315 | $0.63 | $0.60 plus cache reads |
So for the same tokens, introductory Argon equals GPT-6.1 Sol and is half of Opus 5.5. After the introductory period, Argon equals Opus on fresh input and output. Google has not said when that period ends. It has also not said whether the 95% cache discount continues afterward. VentureBeat's $0.20 figure for later cached input is its own stated assumption.
Measured cost per task
Real runs never use identical tokens. Artificial Analysis runs each model through its Intelligence Index v4.3.2, a set of 10 evaluations, and reports what the whole run cost per task. Argon is published at "high" effort only, so high is the one setting where all three can be compared directly.
| Effort | Gemini 4 Argon | Claude Opus 5.5 | GPT-6.1 Sol |
|---|---|---|---|
| Low | Not published | 42 / $0.55 | 42 / $0.13 |
| Medium | Not published | 51 / $1.34 | 48 / $0.21 |
| High | 53 / $1.99 | 54 / $1.82 | 50 / $0.32 |
| xhigh | Not published | 56 / $3.46 | 51 / $0.39 |
| Max | Not published | 58 / $5.98 | 52 / $0.72 |
Each cell is index score / cost per index task. Argon is costed at the $2 / $10 introductory rate. Argon figures are as of October 1, 2026, and the Opus and Sol figures are as of September 30.

Three things follow from the high row:
- Argon costs 6.2 times what GPT-6.1 Sol does per task ($1.99 ÷ $0.32) at the same list price. The gap is tokens consumed. Argon produced 110 million output tokens across the index.
- Argon costs slightly more per task than Opus 5.5 ($1.99 against $1.82), even though Opus charges twice as much per token.
- GPT-6.1 Sol at max effort scores 52 for $0.72. That is one point below Argon at 36% of Argon's cost ($0.72 ÷ $1.99).
Artificial Analysis puts the uncertainty on the composite score at about one point. A 53 against a 54 is a tie, and the spread from 50 to 53 is small.
What the later price would do
If Argon's token use stayed the same and every token were billed at the doubled $4 / $20 rate, its cost per index task would be about $4 (2 × $1.99). That would put it above Opus at xhigh effort ($3.46) for a score three points lower. This is an estimate, since the later cache rate is unpublished.
Vals AI already prices Argon at $4 / $20, and its result points the other way on Opus. On the Vals Index, Argon at high effort cost $15.68 per test against $32.14 for Opus 5.5. According to Kingy AI's notes on the leaderboard, the Opus run used max effort with fallbacks, so the two rows are not at matched effort. Vals also reports that Argon's cost climbs on long agentic tasks, to $193.78 per test on CUA-bench and $57.82 on Code Migration.
The two harnesses agree on one point. Argon's bill depends on how many tokens your task makes it spend, and the token price is a poor predictor of it.
Is Argon actually better than Opus 5.5?
Google's table
Google's launch table compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. Opus has a score on 18 of its rows, and Argon is ahead on 14 of them. Opus leads on four.
| Benchmark in Google's table | Gemini 4 Argon | Claude Opus 5.5 |
|---|---|---|
| Terminal-bench 4.0 | 57.4 | 66.4 |
| FrontierSWE v2 | 55.0 | 62.3 |
| Terminal-Bench Science 0.1 | 57.6 | 63.3 |
| PostTrainBench | 45.3 | 49.3 |
| DeepSWE v1.1 | 77.9 | 74.2 |
| Vibe Code Bench | 91.9 | 90.3 |
| AutomationBench | 51.3 | 42.5 |
| Vals Finance Agent v2 | 65.4 | 58.6 |
| Harvey's Legal Agent Benchmark | 19.6 | 3.8 |
| GraphWalks 256K to 1M | 84.2 | 66.8 |
| LVBench | 91.7 | 83.7 |
The pattern is consistent. Opus leads where an agent works in a terminal for a long time: terminal coding, FrontierSWE, the science terminal suite and ML post-training. Argon's largest margins are in knowledge work, long-context retrieval and video understanding. Some of its other leads are thin, such as CWE-bench v1 at 68.0 against 67.0.

These are vendor-reported numbers, and the table as a whole has not been independently reproduced. Kingy AI's reading of Google's methodology document reports harness differences that affect how far individual rows can be compared: highest thinking settings, a mini-swe harness for DeepSWE, six times the verifier timeout on the science run, and unequal video frame budgets on LVBench. Counting wins also depends on which columns you include. VentureBeat counts 12 outright leads for Argon across all four models, and the count against Opus alone is 14.
Independent measurement
Two third parties have run Argon themselves.
Artificial Analysis has the two models tied at high effort, 53 for Argon and 54 for Opus. Claims that Opus wins 58 to 53 set Opus at max effort against Argon at high. That comparison is fair only if you would pay $5.98 per task for it.
Vals AI ranks Argon first of 41 models on its Vals Index at 68.90%, ahead of Claude Sonnet 5.5 at 67.04% and Opus 5.5 at 66.97%. The margin over Opus is about two points. Vals also flags a weak spot in computer use, where Argon scored 4.83% on CUA-bench, seventh of eight.
Kingy AI reports that Arena's text leaderboard on September 30 had Argon at high effort first with 1,525 against 1,504 for Opus 5.5 at high. Arena measures which answer people prefer, which is a different thing from whether a task was completed correctly.
Taken together: Argon is level with Opus 5.5 at the same effort, a little ahead on finance, legal and long-context work, and behind on terminal-heavy coding agents by Google's own numbers. Nothing here supports calling either one clearly better across the board.
Where GPT-6.1 Sol stands
Google's table has no GPT-6.1 Sol column, so every figure for Sol comes from third parties.
On Artificial Analysis, Sol trails by three to four points at high effort (50 against 53 and 54) for roughly one-sixth of the cost. On the Vals Index, Kingy AI's retrieval of the leaderboard shows GPT-6.1 Sol at 61.15% for $3.24 per test. That is 7.75 points below Argon at about one-fifth of Argon's $15.68.
Sol is the budget choice with a real quality gap, and the size of the gap depends on the suite. It is small on Artificial Analysis' mixed index and larger on Vals' finance, legal and tax work. If the gap matters for your tasks, the choice available today is between Sol and Opus, and Claude Opus 5.5 vs GPT-6.1 Sol works through when Opus pays off. Within OpenAI's lineup, GPT-6.1 Sol vs GPT-6 Astra covers when Astra is worth the step up.
One safety figure circulates with the wrong model attached. Google's prompt-injection chart, as reported by VentureBeat, shows attack success rates of 0.7% for Argon, 1.0% for Opus 5.5 and 27.0% for GPT-6 Sol. That last number is the older Sol. No value for GPT-6.1 Sol has been published.
The 1M output limit
Google says it is "expanding the model's output token limit to an industry-leading 1M tokens, up from the previous 64K tokens." It frames this as headroom to "think deeply and generate hundreds of thousands of tokens in a single trajectory." That describes reasoning plus answer over a long run. It does not promise a million tokens of visible text.
Vals lists the model it evaluated with a 1M-token context window and 262K maximum output. The two figures have not been reconciled, and there is no public API contract to settle it. Opus 5.5 and GPT-6.1 Sol both cap output at 128K.
Plan around 128K-class outputs for now. A larger limit also cuts both ways on cost: at $10 per million, a single response that uses 500K output tokens costs $5, and $10 at the later rate.
What to run now, and what to test when Argon opens
Until Argon has a model ID, the decision is between the two models you can call.
- Start on GPT-6.1 Sol if cost per task drives the decision. Raise effort before you change models. Sol at max scored 52 for $0.72 per index task, which is still below Opus at medium ($1.34).
- Use Opus 5.5 where one failed run costs more than the price difference. Long terminal-agent work is the clearest case, since it is where Opus leads even in Google's table. The Opus 5.5 pricing guide explains what its rates and limit reset change.
- Do not hold a project for Argon. There is no date. Keep the model name in configuration and keep your evaluation tasks saved, so a later swap is a test run and a config change.
On model spend alone, cost per accepted task is cost per attempt divided by pass rate. At the high-effort figures above, Sol stays cheaper than Argon as long as its pass rate on your tasks is more than 16% of Argon's ($0.32 ÷ $1.99). Review time and retries are what move that line, so measure them on your own work.
When Argon reaches paid API customers, check these first:
- Tokens per accepted task on your own workload. This decides the bill. Record input, cache reads, output and reasoning tokens for the same tasks on all three models.
- The end date of the introductory price. Budget at $4 / $20 from the start. If the numbers work only at $2 / $10, they stop working on a date Google has not announced.
- The real output cap, cache write fees, batch discount and any long-prompt tier. None of these is published yet.
- Your terminal and long-running agent tasks. Compare against Opus 5.5 at the same effort, and watch cost on long trajectories, where Vals saw it rise sharply.
- Computer use, if you depend on it. Argon's CUA-bench result on Vals is the weakest number in its profile.
Common questions
Can I use Gemini 4 Argon now?
Not unless you are in Google's Fairwind Program. As of October 1, 2026, Argon has no public API, no model ID and no listing on the Gemini API pricing page. Google says paid API customers and Google AI Ultra subscribers will be first when access widens, without giving a date.
Is Gemini 4 Argon the same price as GPT-6.1 Sol?
Per token, yes during the introductory period: $2 input, $10 output and $0.10 cached input per million, the same as GPT-6.1 Sol for requests up to 272K input tokens. Per task, no. On Artificial Analysis' index at high effort, Argon cost $1.99 per task and GPT-6.1 Sol cost $0.32, because Argon used far more tokens. For other models, the AI API price comparison ranks options by input and output mix.
Is Argon half the price of Opus 5.5?
Only at the introductory token rate and only for identical token counts. On Artificial Analysis' index at high effort, Argon cost slightly more per task than Opus 5.5. After the introductory period, Argon's $4 / $20 rate equals Opus 5.5 on input and output.
Should I wait for Argon or build now?
Build now on Opus 5.5 or GPT-6.1 Sol. Independent results show Argon level with Opus at matched effort, so waiting does not buy a clear step up, and the cost of switching later is small if your evaluation tasks are ready. Earlier leak-based expectations are covered in Gemini 4 vs GPT-6 vs Fable 5.2, which describes what the early demos suggested before the launch.





