# Claude Haiku 5.5 vs GPT-6 Luna: Same Price, 2.7x Cost per Task

> Claude Haiku 5.5 and GPT-6 Luna share $0.10/$0.50 rates under 100K tokens, yet Haiku cost 2.7x more per task at default effort and reprices 5x above 100K.

- URL: https://blog.laozhang.ai/en/posts/claude-haiku-5-5-vs-gpt-6-luna
- Published: 2026-10-08
- Updated: 2026-10-08
- Author: LaoZhang AI Team (https://blog.laozhang.ai/en/about)
- Topic: Model Comparisons
- Tags: Claude Haiku 5.5, GPT-6 Luna, AI Model Comparison, API Pricing, Anthropic, OpenAI

---
Claude Haiku 5.5, released on October 7, 2026, lists the same per-token rates as GPT-6 Luna for ordinary requests: $0.10 per million input tokens, $0.50 per million output tokens and $0.01 per million cached tokens. The parity ends at two different lines. Once a Haiku prompt passes 100,000 tokens, Anthropic bills the whole request at five times those rates. Luna keeps its base rates up to 272,000 input tokens and then charges 2x for input and 1.5x for output.

Equal rates don't produce equal bills. On Artificial Analysis' Intelligence Index tasks at the default `medium` effort, Haiku 5.5 cost $0.047 per task against Luna's $0.017, about 2.7 times as much, because it used more tokens. It also scored higher, 34 against 30. Compare the two at matched scores instead of matched effort settings and Luna still costs roughly 10–30% less per task, while Haiku finishes long tasks in about half the time. Only Haiku reaches scores above 38.

The short answer: start with Luna for high-volume, short, cost-bound work and for any prompt between 100K and 272K tokens. Start with Haiku 5.5 where its higher pass rate, its computer-use results or its speed on long tasks saves more than the extra tokens cost. Prices below are USD per million tokens on each vendor's standard API tier, as of October 8, 2026.

## Is Luna cheaper than Haiku? Same rates up to 100K, then Haiku jumps 5x

Per token, neither is cheaper as long as the prompt stays at or under 100,000 tokens. Above that line, Luna is cheaper on every billing column. Both vendors publish the rates on their pricing pages: [Anthropic's Claude API pricing](https://platform.claude.com/docs/en/about-claude/pricing) and [OpenAI's API pricing](https://developers.openai.com/api/docs/pricing).

| Per 1M tokens (USD) | Haiku 5.5, prompt up to 100K | Haiku 5.5, prompt over 100K | GPT-6 Luna, input up to 272K | GPT-6 Luna, input over 272K |
| --- | --- | --- | --- | --- |
| Input | $0.10 | $0.50 | $0.10 | $0.20 |
| Cached input (cache read) | $0.01 | $0.05 | $0.01 | $0.02 |
| Cache write | $0.125 (5-minute), $0.20 (1-hour) | $0.625 (5-minute), $1 (1-hour) | $0.125 | $0.25 |
| Output, thinking included | $0.50 | $2.50 | $0.50 | $0.75 |
| Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 | $0.05 / $0.25 (Batch or Flex) | $0.10 / $0.375 |
| Faster paid tier | None for Haiku 5.5 | None | Fast mode at 2x | Fast mode at 2x |

Three details in that table change real bills:

- **Both thresholds reprice the whole request.** OpenAI states it directly: prompts with more than 272K input tokens are billed at the higher rates "for the full request". Anthropic labels every Haiku 5.5 price column, output included, "for prompts over 100,000 tokens", so a 101K prompt does not pay the higher rate on just the last 1K.
- **The two thresholds measure slightly different things.** Luna's is "input tokens". Anthropic says only that "a prompt of over 100,000 tokens pays higher prices" and does not say whether cache reads and cache writes count toward that 100K. Its prompt-caching docs define total input as cache reads plus cache writes plus uncached input, which suggests cached tokens count, but no page confirms it. Budget as if they count until Anthropic says otherwise.
- **Output includes thinking.** Haiku 5.5 thinks by default at `medium` effort and Luna reasons by default at `medium`. Those hidden tokens are billed as output on both, and they are one reason the per-task costs below differ so much.

![Claude Haiku 5.5 switches the whole request to 5x rates above 100K prompt tokens, while GPT-6 Luna keeps base rates to 272K input tokens and then charges 2x input and 1.5x output, with three worked requests](https://blog.laozhang.ai/posts/en/claude-haiku-5-5-vs-gpt-6-luna/img/price-thresholds-100k-272k.webp)

OpenAI's Fast mode is what used to be called Priority processing; it was renamed on July 30, 2026. Anthropic's [service tiers page](https://platform.claude.com/docs/en/api/service-tiers) lists Haiku 5.5 among the models without Priority Tier support, and Anthropic's Fast mode covers only its Opus models. If you need a paid lane for lower latency, only Luna has one.

## Same price, 12x, 75% less: what each Haiku 5.5 cost claim measures

All three numbers circulating since launch can be true at once, because each one measures something different.

| Claim | What it actually compares | Does it describe your Haiku vs Luna bill? |
| --- | --- | --- |
| "Haiku 5.5 matches Luna's price" | Per-token list rates for Haiku prompts up to 100K tokens | Yes, for the rate card. It says nothing about how many tokens each model uses. |
| "Haiku cost 12x more than Luna for the same task" | A single-prompt comparison posted on r/ClaudeAI (a voxel pagoda build). On Artificial Analysis' data, about 12x appears only when Haiku at `max` effort is set against Luna at `medium` | No. At the same effort setting, the measured gap is 2.7–5.4x. |
| "Haiku 5.5 costs 75% less" | Anthropic's estimate against **Haiku 4.5**, not against Luna | No. It tells Haiku 4.5 users what their bill may do after moving. |

The 12x figure is easy to reproduce by accident. Artificial Analysis lists Haiku 5.5 at `max` effort at $0.2128 per task and Luna at `medium` at $0.0175: 0.2128 ÷ 0.0175 ≈ 12.2. Haiku at `max` also produced about 162,000 output tokens per task against Luna's 11,456 at `medium`, roughly 14 times as many. Run both models with their defaults or with the same effort level and the ratio falls to the range in the next section. A single prompt can land anywhere, so treat a one-off build as an anecdote rather than a rate.

The 75% figure comes from the [Claude Haiku 5.5 announcement](https://www.anthropic.com/claude-haiku-5-5). Its footnote explains the blend: Haiku 5.5 is priced 90% below Haiku 4.5 for requests up to 100K tokens and 50% below it for longer requests, about 90% of Haiku 4.5 requests were in the shorter group, and the estimate already accounts for Haiku 5.5's newer tokenizer using more tokens for the same text. It is an average for Anthropic's previous traffic mix, not a per-request guarantee.

## Cost per task by effort: Haiku 5.5 spends 2.7–5.4x more than Luna

[Artificial Analysis' Haiku 5.5 vs GPT-6 Luna comparison](https://artificialanalysis.ai/models/releases/comparisons/claude-haiku-5-5-vs-gpt-6-luna) ran both models through its Intelligence Index v4.3.2 (ten evaluations) at every effort level. Because the list rates are identical, the cost differences come almost entirely from token volume.

| Effort | Haiku 5.5 index score | Luna index score | Haiku 5.5 cost per task | Luna cost per task | Haiku ÷ Luna |
| --- | --- | --- | --- | --- | --- |
| `max` | 43 | 38 | $0.2128 | $0.0678 | 3.1x |
| `xhigh` | 41 | 35 | $0.1238 | $0.0422 | 2.9x |
| `high` | 38 | 33 | $0.0794 | $0.0290 | 2.7x |
| `medium` (default on both) | 34 | 30 | $0.0470 | $0.0175 | 2.7x |
| `low` | 29 | 22 | $0.0245 | $0.0045 | 5.4x |
| Luna without reasoning (`none`) | — | 18 | — | $0.0113 | — |

Source: Artificial Analysis, as of October 8, 2026. Ratios are cost per task divided directly, for example 0.0470 ÷ 0.0175 = 2.69.

Three things in this table matter more than the headline ratio:

- **Haiku uses more tokens of every kind.** At `medium`, Haiku 5.5 averaged 32,813 output tokens per task against Luna's 11,456. Output is not even the larger part of the bill. About two thirds of each model's per-task cost is input, mostly cache reads and writes from multi-turn agent work: $0.0306 of Haiku's $0.0470, and $0.0117 of Luna's $0.0175. Haiku's longer runs re-read more context.
- **Luna's cheapest setting is `low`, not `none`.** Without reasoning, Luna used more tokens per task, input and output alike (3,758 output tokens against 2,087 at `low`), and cost $0.0113 per task against $0.0045 while scoring lower. Test both before assuming `none` is the budget option.
- **Haiku's figures assume short prompts.** Artificial Analysis prices Haiku 5.5 at the up-to-100K rates for every task. If any of its tasks sent prompts above 100K tokens, Haiku's real cost on those tasks would be higher than listed.

Run the full index once at `max` effort and the totals look like this: $330 for Haiku 5.5 against $122 for Luna, with 435 million output tokens against 144 million. That 2.7x ratio is slightly lower than the 3.1x per-task ratio because the two runs logged different task counts.

## Is GPT Luna better than Haiku? Haiku 5.5 leads on most scores, not all

On most published scores, Haiku 5.5 is the stronger model at the same effort level. On Artificial Analysis' index it scores higher at every level, by 4 to 7 points. It is not ahead everywhere, and the vendor numbers come with conditions.

Anthropic's launch table, run by Anthropic, compares the two directly:

| Benchmark (Anthropic's launch table) | Haiku 5.5 | GPT-6 Luna |
| --- | --- | --- |
| OSWorld 2.1, offline subset (computer use) | 72.4% | 48.9% |
| Terminal-Bench 4.0 (agentic coding) | 39.2% | 16.4% |
| FrontierCode 1.1, Main | 46.4% | 42.4% |
| Chartography, no tools (chart reasoning) | 46.4% | 29.1% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1620 | 1437 |
| AA-Briefcase v1.1 | 1578 | 1336 |

The table does not say which effort level each model ran at, so read it as Anthropic's best-case comparison. Artificial Analysis' own runs point the same way with smaller margins in places: on Terminal-Bench 4.0 at `max`, it measured 32.8% for Haiku and 12.6% for Luna.

Where Luna leads or ties on Artificial Analysis' runs:

- **AutomationBench-AA:** Luna scores 53.2% at `max` against Haiku's 35.4%, and 47.8% against 33.7% at `high`. Luna is ahead at every effort level except `low`.
- **Strategy & Ops index:** Luna at `max` scores 46 against Haiku's 40, and it leads at every level except `low`.
- **Long-context reasoning (AA-LCR v1.1):** Luna is slightly ahead at every level, 83.3% against 82.7% at `max`.
- **SciCode:** the two stay within about two points of each other, with the lead changing by effort level.

A separate third-party run gives a narrower view. Kingy AI's [40-task Haiku 5.5 vs GPT-6 Luna test](https://kingy.ai/blog/claude-haiku-5-5-vs-gpt-6-luna-benchmark/), published October 7, 2026, sent each task twice through the Claude Max and ChatGPT Pro command-line clients, with every input under 20,000 tokens. Haiku passed every strict check on 60 of 80 outputs (75.0%) and Luna on 51 of 80 (63.75%). Most of the gap came from calculations (8 of 10 against 4 of 10) and editing (8 of 10 against 5 of 10). Both models went 10 for 10 on JSON extraction, instruction following and code fixes. Median completion time through those clients was 6.83 seconds for Haiku and 2.92 seconds for Luna. The sample is small and the two clients wrap requests differently, so treat it as a signal about short, bounded tasks, not a ranking.

## At matched scores, Luna costs 10–30% less and Haiku finishes sooner

Comparing effort settings by name flatters Luna, because Haiku at `medium` is doing harder work than Luna at `medium`. A fairer question is what each model costs to reach the same score. Pairing the closest scores on Artificial Analysis' index:

| Score level | Haiku 5.5 setting and cost | Luna setting and cost | Luna's saving per task | Time per task, Haiku vs Luna |
| --- | --- | --- | --- | --- |
| About 30 | `low`, 29.4, $0.0245 | `medium`, 29.9, $0.0175 | 29% | 62 s vs not published |
| About 34.5 | `medium`, 34.5, $0.0470 | `xhigh`, 34.6, $0.0422 | 10% | 134 s vs 242 s |
| About 38 | `high`, 37.8, $0.0794 | `max`, 38.1, $0.0678 | 15% | 194 s vs 394 s |
| 41–43 | `xhigh` $0.1238, `max` $0.2128 | Not reachable | — | — |

Saving = (Haiku cost − Luna cost) ÷ Haiku cost; for the middle row, (0.0470 − 0.0422) ÷ 0.0470 = 10.2%.

This is the comparison that should drive a default. The 2.7x gap at equal settings shrinks to 10–30% at equal scores. In exchange, Haiku finished each task in 44–51% less time at the two levels where both times are published, and it is the only one of the pair that reaches index scores above 38.

![At matched index scores of about 30, 34.5 and 38, GPT-6 Luna costs 29%, 10% and 15% less per task than Claude Haiku 5.5, while Haiku finishes sooner and alone reaches scores of 41 to 43](https://blog.laozhang.ai/posts/en/claude-haiku-5-5-vs-gpt-6-luna/img/matched-score-cost-per-task.webp)

Latency splits the other way for short calls. On the same page, Luna's time to first token was 3.23 seconds at `low` and 0.87 seconds without reasoning, against 11.99 seconds for Haiku at `low`. For a chat reply or a one-shot classification, Luna answers sooner. For a multi-step task that runs for minutes, Haiku tends to finish first at the same quality.

## Estimate your bill: worked Haiku 5.5 and GPT-6 Luna requests

Per-request cost on either model is:

> (uncached input × input rate + cache reads × cache-read rate + cache writes × cache-write rate + output including thinking × output rate) ÷ 1,000,000

Use Haiku's over-100K rates for the whole request when the prompt exceeds 100,000 tokens, and Luna's over-272K rates when input exceeds 272,000 tokens. The examples below use standard-tier rates and leave caching out unless stated.

| Request | Claude Haiku 5.5 | GPT-6 Luna | Ratio |
| --- | --- | --- | --- |
| 10K input + 1K output | (10,000 × 0.10 + 1,000 × 0.50) ÷ 1M = **$0.0015** | Same rates: **$0.0015** | 1x |
| 99K input + 2K output | (99,000 × 0.10 + 2,000 × 0.50) ÷ 1M = **$0.0109** | Same rates: **$0.0109** | 1x |
| 150K input + 2K output | (150,000 × 0.50 + 2,000 × 2.50) ÷ 1M = **$0.080** | (150,000 × 0.10 + 2,000 × 0.50) ÷ 1M = **$0.016** | 5.0x |
| 300K input + 2K output | (300,000 × 0.50 + 2,000 × 2.50) ÷ 1M = **$0.155** | (300,000 × 0.20 + 2,000 × 0.75) ÷ 1M = **$0.0615** | 2.5x |
| 150K prompt: 140K cache reads + 10K new input + 2K output | (140,000 × 0.05 + 10,000 × 0.50 + 2,000 × 2.50) ÷ 1M = **$0.017** if cache reads count toward 100K | (140,000 × 0.01 + 10,000 × 0.10 + 2,000 × 0.50) ÷ 1M = **$0.0034** | 5.0x |

Rates in the formulas are USD per million tokens.

The cached row is the one to watch in agent loops. If Anthropic does not count cache reads toward the 100K threshold, that Haiku request would cost the same $0.0034 as Luna. Since the docs don't say, $0.017 is the safe budget figure. If your prompts sit near 100K, trimming history or tool output below the line keeps Haiku at its base rates.

A high-volume example shows how the base rates scale. One million classification requests, each with 2,000 input tokens and 300 output tokens:

- **Either model, standard:** (2,000 × 0.10 + 300 × 0.50) ÷ 1M = $0.00035 per request, or **$350** per million requests.
- **Either model, batch** (Luna's Flex tier is the same price): **$175**.
- **Claude Haiku 4.5, for comparison:** (2,000 × 1 + 300 × 5) ÷ 1M = $0.0035 per request, or $3,500.

Thinking changes that picture fast. If each response adds 500 thinking tokens, output rises to 800 tokens and the standard cost becomes (2,000 × 0.10 + 800 × 0.50) ÷ 1M = $0.0006 per request, or $600 per million. That is why the effort setting matters more than the rate card for short tasks: set Luna to `none` or `low` and Haiku to `low` (or turn Haiku's thinking off), then measure real output tokens on a few hundred production requests before committing.

Two token-count caveats apply. Haiku 5.5 uses Anthropic's newer tokenizer, so the same text counts as roughly 30% more tokens than on Haiku 4.5; Anthropic says the exact increase depends on the content. Neither vendor publishes how its token counts compare with the other's, so count your own prompts on both models rather than assuming the same text is the same number of tokens.

## Cost per accepted result: when Haiku's higher pass rate pays off

Token cost per attempt is only half the decision. What you actually pay for is a result you can use:

> cost per accepted result = (model cost per attempt + other cost per attempt) ÷ acceptance rate

"Other cost" is whatever each attempt costs you outside the API bill: a human check, a retry in your pipeline, a failed downstream step.

Here is an illustration with assumed numbers. Take Artificial Analysis' default-effort costs ($0.0470 for Haiku, $0.0175 for Luna) and suppose your own evaluation shows Haiku accepted on 80% of attempts and Luna on 65%.

- **No other cost per attempt:** Haiku 0.0470 ÷ 0.80 = $0.059 per accepted result; Luna 0.0175 ÷ 0.65 = $0.027. Luna wins clearly.
- **$0.25 of review per attempt:** Haiku (0.0470 + 0.25) ÷ 0.80 = $0.371; Luna (0.0175 + 0.25) ÷ 0.65 = $0.412. Haiku wins.
- **Break-even:** 0.65 × (0.0470 + R) = 0.80 × (0.0175 + R) gives R ≈ $0.11 per attempt.

The rule that falls out: once each attempt costs you more in human or pipeline time than the break-even point (about $0.11 with these rates), a measurable gap in acceptance rate outweighs Haiku's higher token bill. When outputs flow straight into an automated system and failures are cheap to retry, Luna's lower cost per attempt wins. The acceptance rates in this example are assumptions. Measure yours on your own tasks at the effort level you plan to ship.

## Which to trial first: Haiku 5.5 or GPT-6 Luna by workload

| Workload | Trial first | Why | Starting setting |
| --- | --- | --- | --- |
| Classification, tagging, routing, short extraction at high volume | GPT-6 Luna | Same rates, lowest cost per task at `low`, faster first token | Luna `low` or `none`; challenge with Haiku `low` |
| Prompts between 100K and 272K tokens | GPT-6 Luna | Haiku bills the whole request at 5x; Luna stays at base rates | Luna `medium` |
| Prompts above 272K tokens | GPT-6 Luna | 2.5x cheaper in the 300K example above | Luna `medium` |
| Business workflow automation | GPT-6 Luna | Leads Haiku on AutomationBench-AA at every effort except `low` | Luna `medium` or `high` |
| Computer use and browser use | Claude Haiku 5.5 | 72.4% vs 48.9% on Anthropic's OSWorld offline subset; new computer and browser toolsets | Haiku `medium` |
| Sub-agents in a Claude-based agent stack | Claude Haiku 5.5 | Anthropic positions it as a sub-agent for Opus 5.5 and Sonnet 5.5 on coding work | Haiku `medium` |
| Multi-step tasks where wall-clock time matters | Claude Haiku 5.5 | 44–51% less time per task at matched scores | Haiku `medium` or `high` |
| Quality above Luna's ceiling | Claude Haiku 5.5, or a larger model | Only Haiku reaches index scores of 41–43 | Haiku `xhigh`, then compare with Sonnet 5.5 |
| Chat replies where first-token latency matters | GPT-6 Luna | 3.23 s at `low` vs 11.99 s for Haiku `low` | Luna `low` |

Two cautions on that table. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding of the kind Terminal-Bench measures, so treat Haiku's lead over Luna there as a ranking between small models, not a reason to downgrade a coding agent. And if you are still choosing within OpenAI's lineup, [GPT-6 Sol vs Terra vs Luna vs Astra: Which Model Should You Use?](https://blog.laozhang.ai/en/posts/gpt-6-sol-vs-terra-vs-luna-vs-astra) covers when Luna is the right OpenAI model in the first place.

Routing between the two is often better than picking one. A practical split is to send short, cheap-to-retry requests to Luna, send computer-use steps and long agent sub-tasks to Haiku, and keep any request above 100K tokens off Haiku. If you only need Luna through an OpenAI-compatible endpoint, the [laozhang.ai model list](https://docs.laozhang.ai/en/models) carries `gpt-6-luna` at the same $0.10/$0.50 rates with the same 272K tier; Claude models are temporarily offline there, so it is not a route to Haiku 5.5.

## Switching models: settings that cause 400 errors or change cost

### Moving a prompt to Claude Haiku 5.5

Anthropic's [Haiku 5.5 migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide) lists the breaking changes. The ones that hit code written for Haiku 4.5 or for OpenAI models:

- **Model ID:** `claude-haiku-5-5`. It is a fixed ID with no date suffix and no separate alias.
- **Sampling parameters:** remove `temperature`, `top_p` and `top_k`. Any non-default value returns a 400 error.
- **Assistant prefill:** a final assistant turn returns a 400 error, even with thinking off. End `messages` with a user turn.
- **Thinking budget:** `thinking: {"type": "enabled", "budget_tokens": N}` returns a 400 error. Use adaptive thinking and set `output_config.effort` to `low`, `medium`, `high`, `xhigh` or `max`.
- **Thinking is on by default** and counts toward `max_tokens`. A small `max_tokens` can stop the response after a thinking block and before any text. Thinking can be disabled only at `high` effort or below.
- **Computer use:** on the Claude API and Google Cloud, use the `computer_toolset_20260801` toolset. The older `computer_20250124` tool returns a 400 error.
- **Refusals:** safety classifiers can end a response with `stop_reason: "refusal"`, and there is no server-side fallback. Handle it in code.
- **Token counts:** recount prompts with the model set to `claude-haiku-5-5` instead of reusing Haiku 4.5 counts, then revisit `max_tokens` and cost estimates.

### Moving a prompt to GPT-6 Luna

- **Model ID:** `gpt-6-luna`. Don't confuse it with `gpt-5.6-luna`, which OpenAI still lists at $0.20 input and $1.20 output.
- **Effort levels:** `reasoning.effort` accepts `none`, `low`, `medium` (default), `high`, `xhigh` and `max`. Luna supports `none`; GPT-6.1 Sol does not.
- **Tool calling:** in Chat Completions, Luna supports function calling only with `reasoning_effort: "none"`. Use the Responses API for reasoning with tools.
- **Sampling parameters:** when effort is anything other than `none`, remove `temperature`, `top_p` and `top_logprobs`.
- **Tiers:** Batch and Flex cost 50% of Standard; Fast mode costs 2x and isn't available with EU data residency for Luna. Regional processing adds 10%.

These details come from [OpenAI's GPT-6 Luna model page](https://developers.openai.com/api/docs/models/gpt-6-luna) and its latest-model migration guide.

## Haiku 5.5 in Claude Code and Claude.ai, Luna in ChatGPT Free and Go

If you use the apps rather than the API, the small model you get changed this week on both sides.

- **Claude.ai:** Free, Pro, Max, Team and Enterprise users can select Haiku 5.5 on web, iOS and Android.
- **Claude Code:** the `haiku` alias resolves to Haiku 5.5 on the Anthropic API and needs Claude Code v2.1.293 or later; run `claude update` if you're behind. On Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, the same alias still resolves to Haiku 4.5, according to [Claude Code's model configuration docs](https://code.claude.com/docs/en/model-config).
- **ChatGPT Chat tab:** OpenAI's October 7 announcement says GPT-6 in Chat is powered by GPT-6 Luna for Free and Go users, with the rollout to those tiers starting October 8, 2026, and by GPT-6 Sol for paid tiers. OpenAI describes both as "tuned for everyday conversation", and its safety page calls the Chat version "GPT-6 Luna (October)", so don't assume Chat behaves exactly like the API model.
- **ChatGPT Work and Codex:** OpenAI says the models there are not changing with this release. Luna is already available in Work and Codex, including for Free and Go users in the desktop app at standard speed, subject to rollout.

App plans bill by subscription rather than by token, so what you see in Claude.ai or ChatGPT says little about API cost. For that, use the per-task numbers above.

## FAQ: Claude Haiku 5.5 vs GPT-6 Luna

### Do cached tokens count toward Claude Haiku 5.5's 100K price threshold?

Anthropic's docs don't say. The pricing page only states that "a prompt of over 100,000 tokens pays higher prices". Its caching docs define total input as cache reads plus cache writes plus uncached input, which points toward cached tokens counting. Until Anthropic clarifies, budget long cached prompts at the over-100K rates and check the bill on a small batch first.

### Should Haiku 4.5 users move to Haiku 5.5 or to GPT-6 Luna?

For prompts under 100K tokens, either move cuts list rates by 90%: the classification example above drops from $3,500 to $350 per million requests on both, before any difference in token counts. Haiku 5.5 keeps your code on the Claude API, but the migration still involves the 400-error changes listed above. Luna wins on cost per task and on prompts above 100K. Anthropic lists Haiku 4.5's retirement as not sooner than October 15, 2026, so plan the test now.

### Is GPT-6 better than Claude overall?

At the small-model tier, Haiku 5.5 outscores GPT-6 Luna on most published benchmarks while costing more per task, as shown above. At the flagship tier the trade-off is different; see [Claude Opus 5.5 vs GPT-6.1 Sol and GPT-6 Sol: When Opus Pays Off](https://blog.laozhang.ai/en/posts/claude-opus-5-5-vs-gpt-6-sol).

### Is Claude Haiku 5.5 faster than GPT-6 Luna?

It depends on the length of the job. Artificial Analysis measured higher output speed for Haiku at every effort level (137 tokens per second at `medium`, 243 at `max`, against 112–127 for Luna) and shorter time per task at matched scores. Luna returns its first token much sooner on short requests: 3.23 seconds at `low` against 11.99 seconds for Haiku at `low`.

## Sources

External pages this guide links to, in the order they appear. Last updated 2026-10-08.

- [Anthropic's Claude API pricing](https://platform.claude.com/docs/en/about-claude/pricing) (platform.claude.com)
- [OpenAI's API pricing](https://developers.openai.com/api/docs/pricing) (developers.openai.com)
- [service tiers page](https://platform.claude.com/docs/en/api/service-tiers) (platform.claude.com)
- [Claude Haiku 5.5 announcement](https://www.anthropic.com/claude-haiku-5-5) (anthropic.com)
- [Artificial Analysis' Haiku 5.5 vs GPT-6 Luna comparison](https://artificialanalysis.ai/models/releases/comparisons/claude-haiku-5-5-vs-gpt-6-luna) (artificialanalysis.ai)
- [40-task Haiku 5.5 vs GPT-6 Luna test](https://kingy.ai/blog/claude-haiku-5-5-vs-gpt-6-luna-benchmark/) (kingy.ai)
- [Haiku 5.5 migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide) (platform.claude.com)
- [OpenAI's GPT-6 Luna model page](https://developers.openai.com/api/docs/models/gpt-6-luna) (developers.openai.com)
- [Claude Code's model configuration docs](https://code.claude.com/docs/en/model-config) (code.claude.com)
