# AI API Price Comparison: Cheapest Model by Input and Output Mix

> As of September 26, 2026, GPT-6 Luna and DeepSeek Flash cost least per token. Among $2-input models, Grok 4.7 is cheapest; past 272K tokens, Sonnet 5 beats Sol.

- URL: https://blog.laozhang.ai/en/posts/cheapest-llm-models
- Published: 2026-07-02
- Updated: 2026-09-26
- Author: LaoZhang AI Team (https://blog.laozhang.ai/en/about)
- Category: API Pricing
- Tags: LLM API, AI Model Pricing, OpenAI API, Claude API, Gemini API, Grok API, DeepSeek API

---
As of September 26, 2026, the cheapest major AI API per token is OpenAI's GPT-6 Luna at $0.10 per 1M input tokens and $0.50 per 1M output tokens. DeepSeek Flash is close behind at $0.15 and $0.60 during off-peak hours, and it has the lowest cached-input rate of any model here. One price band up, GPT-6 Sol and Claude Sonnet 5 both list at $2 input and $10 output, while Grok 4.7 charges the same $2 for input but only $6 for output.

Those headline numbers rarely decide the bill on their own. Four things do: how much of your traffic is output, how much of your input hits a cache, whether single prompts run past 200K or 272K tokens, and whether your budget extends past December 31, 2026, when Gemini 3.8 Flash's introductory price ends. Use the table below to pick where to start, then check the condition in the right-hand column against your own traffic.

| Your workload | Lowest list price to test first | What would change the answer |
|---|---|---|
| High-volume short answers, extraction, classification | GPT-6 Luna ($0.10 / $0.50) or DeepSeek Flash off-peak ($0.15 / $0.60) | Quality misses your bar; DeepSeek's peak-hour rate is double |
| The same long instructions or documents reused on every call | DeepSeek Flash ($0.003 cached input off-peak) or GPT-6 Luna ($0.01 cached input) | Your real cache-hit share is low |
| A mid-price general model | Grok 4.7 ($2 / $6), then Sol or Sonnet 5 ($2 / $10) | Prompts of 200K tokens or more, which Grok 4.7 bills at $4 / $12 for the whole request |
| Prompts above 272K tokens | Claude Sonnet 5, which keeps $2 / $10 up to its 1M-token context | Sol reprices the whole request at $4 / $15 above 272K |
| Jobs that can wait minutes to hours | Batch or Flex on OpenAI, Batch on Claude and Gemini, all at half price | You need an immediate response |

## Current AI API prices per 1M tokens

All prices below are first-party list prices in US dollars per 1 million tokens, Standard processing, short context, as of September 26, 2026. They come from the model makers' own pages: [OpenAI API pricing](https://developers.openai.com/api/docs/pricing), [Claude API pricing](https://platform.claude.com/docs/en/about-claude/pricing), [Gemini Developer API pricing](https://ai.google.dev/gemini-api/docs/pricing), [xAI API pricing](https://docs.x.ai/developers/pricing) and [DeepSeek API pricing](https://api-docs.deepseek.com/quick_start/pricing). Gemini rows apply to an AI Studio key, not Vertex AI. A dash means the vendor page shows no rate for that column.

The rows are grouped by output price, because output is usually the larger part of the bill. The groups are price bands, not quality rankings.

### Low-cost models: output at $5 or less

| Model | Input | Cached input | Output | What changes this price |
|---|---:|---:|---:|---|
| GPT-6 Luna (`gpt-6-luna`) | $0.10 | $0.01 | $0.50 | Prompts over 272K tokens bill the whole request at $0.20 / $0.75 |
| DeepSeek Flash (`deepseek-flash`) | $0.15 off-peak, $0.30 peak | $0.003 / $0.006 | $0.60 / $1.20 | Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday |
| Gemini 3.1 Flash-Lite | $0.25 (audio $0.50) | $0.025 | $1.50 | Free tier exists, but free-tier content is used to improve Google products |
| Gemini 3.5 Flash-Lite | $0.30 | – | $2.50 | Batch: $0.15 / $1.25 |
| DeepSeek V4-Pro (`deepseek-v4-pro`) | $0.66 off-peak, $1.32 peak | $0.022 / $0.044 | $1.98 / $3.96 | Same peak schedule as Flash |
| Gemini 3.8 Flash | $0.75 | $0.075 plus storage | $3.75 | Introductory rate through December 31, 2026; $1.50 / $7.50 from January 1, 2027 |
| Claude Haiku 4.5 | $1 | $0.10 | $5 | Batch: $0.50 / $2.50 |

### Mid-price models: output from $6 to $12

| Model | Input | Cached input | Output | What changes this price |
|---|---:|---:|---:|---|
| Grok 4.7 (`grok-4.7`) | $2 | $0.50 | $6 | Prompts of 200K tokens or more bill the whole request at $4 / $1 cached / $12; no Batch |
| GPT-6 Sol (`gpt-6-sol`) | $2 | $0.20 | $10 | Prompts over 272K tokens bill the whole request at $4 / $15 |
| Claude Sonnet 5 | $2 | $0.20 | $10 | Same rate across the full 1M-token context |
| Gemini 3.1 Pro Preview | $2 | – | $12 | Prompts over 200K tokens: $4 / $18 |

### Premium models: output at $20 or more

| Model | Input | Cached input | Output | What changes this price |
|---|---:|---:|---:|---|
| Claude Opus 5.5 | $4 | $0.20 | $20 | Fast mode (research preview): $8 / $40 |
| GPT-6 Astra (`gpt-6-astra`) | $10 | $1 | $50 | Prompts over 272K tokens: $20 / $75 for the whole request |
| Claude Fable 5.1 | $10 | $0.25 | $50 | Cache hits cost only 2.5% of the input rate |

Two notes keep the cached column honest. OpenAI now lists a separate cache-write rate for GPT-6 models ($2.50 per 1M for Sol), and Claude charges 1.25 times the input rate to write a 5-minute cache or 2 times for a 1-hour cache. Cached-input savings only show up once the same prefix is read back more often than it is rewritten. Gemini also bills cache storage by the hour: $0.50 per 1M tokens per hour on Gemini 3.8 Flash through 2026, then $1.00.

## The billing rules that decide which model is cheaper

A price table ranks models by one pair of numbers. Your invoice depends on how your traffic meets each vendor's rules, and the rules differ more than the rates do.

### Output share

Grok 4.7, Sol and Sonnet 5 all charge $2 per 1M input tokens, so for the same token counts on uncached, short-context traffic, Grok 4.7's list price can never be higher than theirs. How much it saves depends on output share: about 13% on an input-heavy month and about 36% on an output-heavy one, as the examples below show. The same logic separates DeepSeek Flash at peak rates from Gemini 3.1 Flash-Lite: Flash's peak input costs more ($0.30 vs. $0.25), but its output costs less ($1.20 vs. $1.50), so the order flips once output takes a large enough share.

### Cache hits

When most of your input is a repeated system prompt, tool list or reference document, the cached rate matters more than the input rate. Cached input costs a tenth of the normal input rate on GPT-6 and Claude Haiku 4.5 and Sonnet 5, 5% on Claude Opus 5.5, and 2.5% on Claude Fable 5.1. DeepSeek Flash's cached input costs 2% of its cache-miss rate, and V4-Pro's about 3%. Grok 4.7's cached input is a quarter of its input rate ($0.50 vs. $2), so caching narrows its lead over Sol and Sonnet 5. The side effect is that a cache-heavy workload shrinks the input side of every bill until output price is almost the only thing left to compare. For the details of how writes and reads are billed, see [OpenAI vs Claude cache pricing](https://blog.laozhang.ai/en/posts/openai-vs-claude-cache-pricing).

### Long prompts: the 200K and 272K thresholds

This is where two models with identical list prices stop being equal. OpenAI prices any GPT-6 prompt with more than 272K input tokens at 2 times the input and cache rates and 1.5 times the output rate, and the multiplier applies to the full request, not only the tokens above 272K. Gemini 3.1 Pro Preview switches to $4 / $18 above 200K. Grok 4.7's threshold is also 200K: once a prompt reaches it, every token in the request bills at $4 input, $1 cached input and $12 output. Claude models from the 4.6 generation onward keep the standard rate across the whole 1M-token window, so a 900K-token request is billed at the same per-token rate as a 9K one. DeepSeek lists one rate for its 1M context.

### Gemini 3.8 Flash doubles on January 1, 2027

Gemini 3.8 Flash's $0.75 / $3.75 is an introductory price with a published end date. From January 1, 2027, it becomes $1.50 / $7.50, and cached input moves from $0.075 to $0.15. If you are signing up for a year of traffic, budget with the 2027 rate. At that rate, 3.8 Flash costs more than Grok 4.7 on output-heavy work.

### DeepSeek peak hours

DeepSeek charges double between 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, except on Chinese public holidays; weekends are off-peak all day. In US time zones those windows fall in the evening and overnight, roughly 9 p.m. to midnight and 2 a.m. to 6 a.m. Eastern while daylight saving time is in effect, so traffic during the US workday is billed at the off-peak rate. Evening and overnight jobs run on a US schedule are the ones that hit peak pricing. [DeepSeek V4 peak pricing](https://blog.laozhang.ai/en/posts/deepseek-v4) walks through rescheduling math.

### Batch and Flex: half price if you can wait

OpenAI's Batch and Flex tiers are exactly half the Standard rate on all three GPT-6 models, and Claude and Gemini Batch are also 50% off both input and output. Batch is asynchronous: you submit a file of requests and collect results later, so it fits nightly classification or evaluation runs, not chat. Flex trades latency and availability for the lower price. Grok 4.7 does not support xAI's Batch API. DeepSeek's pricing page shows no Batch rate; its discount is the off-peak window. Going the other way, OpenAI's Fast mode (renamed from Priority on July 30, 2026) and xAI's Priority processing both cost twice the standard rate.

### Token counts are not the same across vendors

A dollar per 1M tokens only compares cleanly if the same text becomes the same number of tokens. Anthropic states that Claude 4.7 and later models, which include Sonnet 5, Opus 5.5 and Fable 5.1, use a newer tokenizer that produces about 30% more tokens for the same text than Claude's earlier tokenizer, with the exact increase depending on content. No vendor publishes a cross-vendor ratio, so the only reliable number is the token count your own requests report in each API's usage field.

### Data residency adds 10%

If you need in-region processing, OpenAI's regional data residency endpoints add 10% for eligible models released on or after March 5, 2026, Claude's US-only inference option multiplies all token prices by 1.1, and xAI's US regional endpoint bills Grok tokens at 1.1 times the global rate. Default global processing is billed at the standard rate.

## What a month costs: worked examples

The monthly formula is the same for every vendor:

```text
monthly cost = (uncached input M × input rate)
             + (cached input M × cached rate)
             + (output M × output rate)

M = millions of tokens per month; rates in $ per 1M tokens
```

The two examples below use Standard processing, short context and no caching. Example 1 is an input-heavy month such as a support bot or document Q&A: 300M input and 30M output tokens. Example 2 is output-heavy, such as code or long-form drafting: 20M input and 40M output tokens. These are estimates from list prices, not measured invoices. They exclude retries, tool fees, cache-write charges and tokenizer differences.

| Model | Example 1: 300M in / 30M out | Example 2: 20M in / 40M out |
|---|---:|---:|
| GPT-6 Luna | $45 | $22 |
| DeepSeek Flash, off-peak | $63 | $27 |
| Gemini 3.1 Flash-Lite | $120 | $65 |
| DeepSeek Flash, peak | $126 | $54 |
| Gemini 3.5 Flash-Lite | $165 | $106 |
| DeepSeek V4-Pro, off-peak | $257.40 | $92.40 |
| Gemini 3.8 Flash, 2026 price | $337.50 | $165 |
| Claude Haiku 4.5 | $450 | $220 |
| DeepSeek V4-Pro, peak | $514.80 | $184.80 |
| Gemini 3.8 Flash, 2027 price | $675 | $330 |
| Grok 4.7 | $780 | $280 |
| GPT-6 Sol | $900 | $440 |
| Claude Sonnet 5 | $900 | $440 |
| Gemini 3.1 Pro Preview | $960 | $520 |
| Claude Opus 5.5 | $1,800 | $880 |
| GPT-6 Astra | $4,500 | $2,200 |
| Claude Fable 5.1 | $4,500 | $2,200 |

To check one row yourself: Grok 4.7 in Example 1 is 300 × $2 + 30 × $6 = $600 + $180 = $780.

The order shifts between the two columns. DeepSeek Flash at peak rates costs more than Gemini 3.1 Flash-Lite in Example 1 but less in Example 2, because its output rate is lower. DeepSeek V4-Pro off-peak passes Gemini 3.5 Flash-Lite once output dominates. Gemini 3.8 Flash at its 2027 price stays below Grok 4.7 for input-heavy work but ends up above it for output-heavy work.

![Three model pairs that swap places between the input-heavy month and the output-heavy month, such as DeepSeek Flash at peak costing $126 vs $120 for Gemini 3.1 Flash-Lite, then $54 vs $65](https://blog.laozhang.ai/posts/en/cheapest-llm-models/img/monthly-cost-order-flip.webp)

### The same input-heavy month with 80% cache hits

Now assume 240M of Example 1's 300M input tokens are cache hits. Cache-write charges are left out.

| Model | Without cache | With 80% cache hits |
|---|---:|---:|
| GPT-6 Luna | $45 | $23.40 |
| DeepSeek Flash, off-peak | $63 | $27.72 |
| DeepSeek Flash, peak | $126 | $55.44 |
| Gemini 3.8 Flash, 2026 price | $337.50 | $175.50 plus hourly storage |
| Claude Haiku 4.5 | $450 | $234 |
| Grok 4.7 | $780 | $420 |
| GPT-6 Sol | $900 | $468 |
| Claude Sonnet 5 | $900 | $468 |
| Claude Opus 5.5 | $1,800 | $888 |

Sol and Sonnet 5 drop by 48%, and Claude Opus 5.5 with a good cache comes in below uncached Sol. Grok 4.7 drops by 46% (60 × $2 + 240 × $0.50 + 30 × $6 = $420). It stays below cached Sol and Sonnet 5, but its lead shrinks from $120 to $48 because its cached rate is $0.50 rather than $0.20.

### One long request: 400K input, 5K output

Sol and Sonnet 5 have the same short-context price. A single 400K-token prompt with a 5K-token answer separates them:

- GPT-6 Sol, above 272K so the whole request reprices: 0.4 × $4 + 0.005 × $15 = $1.675
- Claude Sonnet 5, standard rate across 1M: 0.4 × $2 + 0.005 × $10 = $0.85
- Gemini 3.1 Pro Preview, above 200K: 0.4 × $4 + 0.005 × $18 = $1.69
- Grok 4.7, at or above 200K so the whole request reprices: 0.4 × $4 + 0.005 × $12 = $1.66

A naive calculation at Sol's short-context rate would also give $0.85, which is how long-context pipelines end up with bills nearly double the estimate.

![Rates by prompt size for GPT-6 Sol, Claude Sonnet 5 and Gemini 3.1 Pro Preview, and the cost of one 400K-token request: $1.675, $0.85 and $1.69](https://blog.laozhang.ai/posts/en/cheapest-llm-models/img/long-prompt-400k-request.webp)

## Estimate your own bill

Measure before you estimate. Send 50 to 100 real requests from your application to each candidate and read the input, cached-input and output token counts from the usage field of each response. Average them, multiply by expected monthly volume, and apply the formula. The short script below does that arithmetic with the rates above; replace the numbers with your own volumes and recheck the rates on the vendor pages before you commit spend.

```python
# USD per 1M tokens: (input, cached input, output)
# Standard tier, short context, list prices as of September 26, 2026
# Use None for a model with no published cached-input rate
PRICES = {
    "GPT-6 Luna":              (0.10, 0.01, 0.50),
    "GPT-6 Sol":               (2.00, 0.20, 10.00),
    "Claude Haiku 4.5":        (1.00, 0.10, 5.00),
    "Claude Sonnet 5":         (2.00, 0.20, 10.00),
    "Claude Opus 5.5":         (4.00, 0.20, 20.00),
    "Gemini 3.1 Flash-Lite":   (0.25, 0.025, 1.50),
    "Gemini 3.8 Flash":        (0.75, 0.075, 3.75),  # doubles on Jan 1, 2027
    "Grok 4.7":                (2.00, 0.50, 6.00),
    "DeepSeek Flash off-peak": (0.15, 0.003, 0.60),
    "DeepSeek Flash peak":     (0.30, 0.006, 1.20),
}

def monthly_cost(model, input_m, output_m, cache_hit_share=0.0):
    """input_m and output_m are millions of tokens per month."""
    rate_in, rate_cached, rate_out = PRICES[model]
    cached_m = input_m * cache_hit_share
    if cached_m and rate_cached is None:
        raise ValueError(f"{model}: no published cached-input rate")
    cached_cost = cached_m * rate_cached if cached_m else 0.0
    return (input_m - cached_m) * rate_in + cached_cost + output_m * rate_out

for model in PRICES:
    try:
        cost = monthly_cost(model, input_m=300, output_m=30, cache_hit_share=0.8)
        print(f"{model:<24} ${cost:,.2f}")
    except ValueError as err:
        print(err)
```

With the values shown, the script reproduces the 80% cache table, prints $66.00 for Gemini 3.1 Flash-Lite and $420.00 for Grok 4.7. Gemini cache storage fees are not included. Set `cache_hit_share=0.0` to get the uncached Example 1 column.

Then add what the formula leaves out: failed calls you retry, tool or search fees, and the 10% residency uplift if it applies. Gemini's Google Search grounding, for example, includes 5,000 free requests a month shared across Gemini 3.x models, then costs $14 per 1,000. For agents that loop on their own, put a hard spending cap in front of the API call; [LLM agent API spend kill switch](https://blog.laozhang.ai/en/posts/llm-agent-api-spend-kill-switch) covers how.

## Free tiers, resellers and price tracker sites

Among the five vendors here, Google lists free-of-charge input and output for Gemini 3.8 Flash and Gemini 3.1 Flash-Lite on its Free tier. The catch is written in the same table: content sent on the Free tier is used to improve Google's products, and content on the paid tier is not. That makes the free tier suitable for trying prompts on public or synthetic data, not for customer records or private code. Rate limits and other vendors' free offers are covered in [Best Free AI API Provider: Pick by Limits, Data Use, and Region](https://blog.laozhang.ai/en/posts/free-ai-api-tiers-compared), and funding a paid Gemini key is covered in [Gemini API key pricing in the US](https://blog.laozhang.ai/en/posts/gemini-api-pricing).

Gateways, aggregators and resellers sell access to the same models at their own prices. Their rows are that company's price, with that company's refund terms, rate limits and support, and they can differ from the vendor page in either direction. Price tracker sites are useful for spotting new models quickly, but a tracker row is only as current as its last update. Before you buy credits on any of them, match the exact model ID and the check date against the vendor's own pricing page. [Cheapest LLM API provider](https://blog.laozhang.ai/en/posts/cheapest-llm-api-provider) compares these options on price, latency and risk.

## How to shortlist two models

Pick two, not ten. Start with the cheapest row that fits your workload in the first table, then add one model from the next price band as a quality check. For most readers that means one of these pairs:

- High-volume text processing: GPT-6 Luna and DeepSeek Flash, then Gemini 3.1 Flash-Lite if you want a third vendor.
- General assistant or coding at mid price: Grok 4.7 and either Sol or Sonnet 5, depending on which SDK and tools you already use.
- Very long documents: Claude Sonnet 5, with Gemini 3.1 Pro Preview as the comparison, keeping its 200K threshold in mind.
- Work where a wrong answer is expensive: Claude Opus 5.5 before GPT-6 Astra or Claude Fable 5.1, since it costs 40% of their list price.

Run the same real requests through both, count how many results you would actually accept, and divide the total cost by that number. A model that costs half as much per token but needs a retry on every third request is not half the price. The lowest cost per accepted result is the one to put into production.

## FAQ

### Is the Claude API or the OpenAI API cheaper?

It depends on the model pair. At the low end, GPT-6 Luna ($0.10 / $0.50) is a tenth of Claude Haiku 4.5's list price ($1 / $5). In the middle, GPT-6 Sol and Claude Sonnet 5 are identical at $2 / $10 until a prompt passes 272K tokens. Above that, Sol's input rate doubles for the whole request, so a long, input-dominated request costs about half as much on Sonnet 5. At the top, Claude Opus 5.5 ($4 / $20) lists at 40% of GPT-6 Astra ($10 / $50), and Claude Fable 5.1 matches Astra. Claude's newer tokenizer counts more tokens for the same text, so compare actual usage numbers from your own requests. [Claude API vs OpenAI API pricing](https://blog.laozhang.ai/en/posts/claude-api-vs-openai-api-pricing) shows the pairwise method in more detail.

### Which AI model is the cheapest per token?

As of September 26, 2026, GPT-6 Luna has the lowest standard input and output rates among the major first-party APIs compared here ($0.10 and $0.50 per 1M tokens). DeepSeek Flash has the lowest cached-input rate ($0.003 off-peak) and is second on standard rates during off-peak hours. Per token is not the same as per task, though: shorter answers or fewer retries can make a pricier model cheaper overall.

### Which AI API has the best free API key?

Google's Gemini Developer API lists a free tier for Gemini 3.8 Flash and Gemini 3.1 Flash-Lite, with free input and output on that tier. The trade-off is that free-tier content is used to improve Google's products. For anything beyond testing with non-sensitive data, a paid key is the safer choice.

### Will Gemini 3.8 Flash get more expensive?

Yes. Google's pricing page shows $0.75 input and $3.75 output through December 31, 2026, and $1.50 and $7.50 starting January 1, 2027. Batch rates and context caching double on the same date.

### Why do AI API price comparison sites show different numbers?

They update on different schedules, and some list older model generations or a reseller's price next to vendor prices. Each vendor's pricing page is the one that matches your invoice, so use trackers to find candidates and the vendor page to confirm the rate.
