# Cheapest LLM API Provider: Where the Same Model Costs Less

> No provider wins for every model. As of September 26, 2026, DeepSeek direct off-peak costs half of hosts charging its peak rate; check versions and router fees.

- URL: https://blog.laozhang.ai/en/posts/cheapest-llm-api-provider
- Published: 2026-07-01
- Updated: 2026-09-26
- Author: LaoZhang AI Team (https://blog.laozhang.ai/en/about)
- Category: API Pricing
- Tags: LLM API, API Pricing, DeepSeek API, OpenRouter, Open-Weight Models

---
No single company sells LLM inference cheapest across the board. The cheapest provider depends on which model you're buying, what hours your traffic runs, and whether two price rows are even the same model version. Prices below are as of September 26, 2026, in USD per 1 million tokens.

Three patterns decide most cases:

- **Closed models** (GPT-6, Claude, Gemini, Grok) are priced by their developer. A router passes that price through and adds its own fee. A reseller sets its own price, which you check model by model.
- **Open-weight models** (DeepSeek, Kimi, Qwen, GLM, gpt-oss) are sold by many hosts, and the winner flips by model. DeepInfra lists DeepSeek V4 Pro and Kimi K3 below Together; Together lists Qwen3.7-Max below DeepInfra.
- **DeepSeek's own API** charges half price outside two weekday UTC windows. Together's DeepSeek rows match DeepSeek's peak price at all hours, so off-peak traffic is cheapest when you buy direct.

| Your situation | Usually cheapest | Check before switching |
| --- | --- | --- |
| You use a closed model | The developer's own API, with Batch or Flex for work that can wait | Whether a router fee or reseller multiplier buys you something you need |
| You use DeepSeek and most calls land outside 01:00–04:00 and 06:00–10:00 UTC on weekdays | DeepSeek direct at off-peak rates | Your real hour-by-hour traffic mix |
| You need DeepSeek V4 Pro during peak hours | DeepInfra ($1.30 input / $2.60 output) over Together or DeepSeek peak ($1.32 / $3.96) | That DeepInfra's unlabeled V4 Pro snapshot matches the one you tested |
| You use another open-weight model | Whichever host lists that exact model lower; it varies | Snapshot date, context length, and precision label |
| You switch models often or want automatic fallback | A router such as OpenRouter | A 5.5% credit-purchase fee on the Standard plan |

Which model to buy is a separate decision. [AI API Price Comparison: Cheapest Model by Input and Output Mix](https://blog.laozhang.ai/en/posts/cheapest-llm-models) covers that. Everything below assumes you already know the model and want the cheapest place to call it.

## Four kinds of LLM API providers

"Provider" covers four different businesses. They sell different things, so their prices mean different things.

| Type | Examples | What you're paying for | Where the price comes from | What to watch |
| --- | --- | --- | --- | --- |
| First-party API | OpenAI, Anthropic, Google, xAI, DeepSeek | The model from the company that trains it, including its batch, flex, or off-peak discounts | The developer's pricing page | Date-bound prices and retired model names |
| Open-model host | Together, DeepInfra, Fireworks, Groq, SiliconFlow | Open weights served on the host's hardware | The host's own table, set model by model | Snapshot version, precision, and billing unit |
| Router | OpenRouter | One API and one balance across many upstream providers | The upstream provider's price, plus a platform fee | Credit-purchase fees, BYOK fees, plan limits |
| Gateway or reseller | laozhang.ai, regional resellers | One compatible endpoint, prepaid balance, often local payment | The gateway's own table, sometimes with group multipliers | Which model actually answers, and who supports you when upstream fails |

The first-party row is the reference for closed models. As of September 26, 2026, OpenAI lists GPT-6 Luna at $0.10 input and $0.50 output, with Batch and Flex at half price. Google lists Gemini 3.1 Flash-Lite at $0.25 / $1.50. Anthropic lists Claude Haiku 4.5 at $1 / $5, and xAI lists grok-4.7 at $2 / $6. Open-model hosts don't serve these models. Other sellers, such as cloud platforms, routers, and resellers, either pass the developer's price through or set their own, which you verify on their page. For side-by-side closed-model pricing, see [Claude API vs OpenAI API Pricing: When Each One Costs Less](https://blog.laozhang.ai/en/posts/claude-api-vs-openai-api-pricing).

Open-weight models are where shopping across providers pays. Each host picks which snapshot it runs, at what precision, and at what margin. Fireworks bills serverless models per token, postpaid, with separate Standard, Priority, and Fast tiers, so even one host can list several prices for one model. DeepInfra says some of its language models are priced per token and most of its other models are billed by execution time.

## Same model, different monthly bill

The table prices one month of 100 million input tokens and 20 million output tokens with no caching. Each figure is `input millions × input rate + output millions × output rate`, so you can redo it with your own volumes.

| Model | Provider and row | Input / output per 1M | Month (100M in, 20M out) |
| --- | --- | ---: | ---: |
| DeepSeek V4 Pro | DeepSeek direct, off-peak | $0.66 / $1.98 | $105.60 |
| DeepSeek V4 Pro | DeepSeek direct, peak | $1.32 / $3.96 | $211.20 |
| DeepSeek V4 Pro | Together (V4 Pro 0813) | $1.32 / $3.96 | $211.20 |
| DeepSeek V4 Pro | DeepInfra (DeepSeek-V4-Pro, no date label) | $1.30 / $2.60 | $182.00 |
| DeepSeek V4.1 Flash | DeepSeek direct, off-peak | $0.15 / $0.60 | $27.00 |
| DeepSeek V4.1 Flash | DeepSeek direct, peak | $0.30 / $1.20 | $54.00 |
| DeepSeek V4.1 Flash | Together | $0.30 / $1.20 | $54.00 |
| Kimi K3 | Together | $3.00 / $15.00 | $600.00 |
| Kimi K3 | DeepInfra | $2.85 / $14.25 | $570.00 |
| Qwen3.7-Max | Together | $1.50 / $4.50 | $240.00 |
| Qwen3.7-Max | DeepInfra | $2.50 / $7.50 | $400.00 |

Prices from [DeepSeek pricing](https://api-docs.deepseek.com/quick_start/pricing), [Together pricing](https://www.together.ai/pricing), and [DeepInfra pricing](https://deepinfra.com/pricing).

Three conclusions hold for this mix:

- **Neither host is cheaper overall.** DeepInfra saves $30 a month on Kimi K3 and $29.20 on V4 Pro. Together saves $160 on Qwen3.7-Max.
- **Time of day beats host choice for DeepSeek.** Off-peak direct V4 Pro costs $105.60, well under DeepInfra's $182.00. During peak hours, DeepInfra is lower than both DeepSeek and Together, provided its snapshot passes your tests.
- **Caching widens the gap.** With 60% of input served from cache, a V4.1 Flash month costs $18.18 at DeepSeek off-peak and $36.36 on Together: 40M × $0.15 + 60M × $0.003 + 20M × $0.60, against 40M × $0.30 + 60M × $0.006 + 20M × $1.20.

![Bar chart of monthly DeepSeek bills for 100M input and 20M output tokens: V4 Pro costs $105.60 direct off-peak, $182.00 on DeepInfra, and $211.20 at DeepSeek peak or on Together](https://blog.laozhang.ai/posts/en/cheapest-llm-api-provider/img/deepseek-monthly-bill.webp)

DeepSeek's peak windows are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays. During US daylight saving time that is 9 PM–midnight and 2–6 AM Eastern. A 9-to-5 workday on either US coast falls entirely outside the peak windows. For a job-by-job scheduling test, see [DeepSeek V4 Peak Pricing: Calculate the Real Cost Before You Reschedule](https://blog.laozhang.ai/en/posts/deepseek-v4).

## Four ways a cheap price row misleads

![Four traps behind a cheap LLM price row: a host charging the peak rate at all hours, an older model snapshot, a router credit fee, and a ranking written by a listed provider](https://blog.laozhang.ai/posts/en/cheapest-llm-api-provider/img/misleading-price-rows.webp)

### The host charges the developer's peak price

Together's DeepSeek V4 Pro 0813 row ($1.32 / $3.96) and V4.1 Flash row ($0.30 / $1.20) are identical to DeepSeek's peak prices. Together's page shows no off-peak discount. A host that matches the peak price looks competitive in a table that only shows DeepSeek's peak row, but it costs double whenever DeepSeek's off-peak price applies. A host can still be worth it for reasons other than price, such as one account for several open models. Just price it against the rate you would actually pay direct.

### The cheaper row is an older snapshot

DeepSeek retired the `deepseek-v4-flash` API name. Calls to it are now served by V4.1-Flash at the Flash price. Hosts still sell the earlier open-weight snapshot: Together lists DeepSeek V4 Flash 0731 at $0.14 / $0.28, and DeepInfra lists DeepSeek-V4-Flash-0731 at $0.06 / $0.18. DeepInfra also has a row simply labeled DeepSeek-V4-Flash at $0.09 / $0.18, and its page doesn't say which snapshot that is.

Those rows are real and cheap. On the month above, the 0731 snapshot costs $19.60 on Together and $9.60 on DeepInfra, versus $27.00 for V4.1 Flash off-peak direct. They are not the model DeepSeek serves today, though. An older snapshot can still win on your workload. Suppose, hypothetically, that 30% of its outputs fail your checks and you rerun those tasks on V4.1 Flash off-peak. The month then costs about $9.60 + 0.30 × $27.00 = $17.70, still below $27.00. The mistake is assuming the rows are interchangeable, not choosing the older snapshot after testing it.

Precision matters in the same way. Some host listings carry labels such as FP8 or FP4 in the model name. Read the model card for the snapshot date, context length, and precision before comparing prices.

### The router adds a fee on top of pass-through prices

OpenRouter says it passes through the underlying provider's price without markup and charges a fee when you buy credits: 5.5% on the Standard plan and 8% on Business, with discounts negotiated for Enterprise. A $500 top-up on Standard carries a $27.50 fee. The V4 Pro month that costs $211.20 at a provider's list price therefore costs about $222.82 through OpenRouter credits.

With your own provider keys (BYOK), Standard and Business include $25,000 of list-price usage per month with no fee, then charge 5%. The Free plan has no platform fee but limits you to 25+ free models and 50 requests per day, with no BYOK. The pass-through price is whichever upstream provider serves the call, so check which provider that is on the model page rather than assuming the lowest row in this table. See [OpenRouter pricing](https://openrouter.ai/pricing) and the [OpenRouter FAQ](https://openrouter.ai/docs/faq).

A router pays for itself when switching models, falling back across providers, or consolidating billing saves more engineering time than the fee costs. It is not a way to get a lower per-token price than the provider it routes to.

### The ranking comes from a provider on the list

"Cheapest provider" lists written by a host that includes itself are marketing. So are router pages quoting a headline price per million tokens without naming the model and version. A useful comparison prices the same model snapshot on each provider, shows input, cached input, and output separately, and dates the prices. If a ranking doesn't do that, use it only as a list of names to check.

## Effective cost: the calculation that decides it

A price table is an input, not the answer. Price a real month of your own traffic, then divide by the work that actually passed.

```text
monthly cost = ( uncached input M × input rate
               + cached input M  × cached rate
               + output M        × output rate ) × (1 + platform fee)

cost per accepted output = monthly cost ÷ outputs that pass your checks
```

Count retries in the token volumes, because every failed call you retry is billed again. Then adjust for the factors that change which provider wins:

| Factor | Why it moves the answer | Example as of September 26, 2026 |
| --- | --- | --- |
| Time or mode discounts | The same model can cost half as much at a different hour or in batch | DeepSeek off-peak is half of peak; OpenAI Batch and Flex and Gemini Batch are half price |
| Cached input | Long, repeated context can dominate agent and RAG bills | V4.1 Flash cached input is $0.003 at DeepSeek off-peak and $0.006 on Together |
| Platform fees | A fee on credits applies to every token you buy | OpenRouter: 5.5% Standard, 8% Business |
| Version and precision | A cheaper row may be a different model | V4 Flash 0731 rows versus DeepSeek's current V4.1-Flash |
| Billing unit | Per token, per call, and per second of execution don't compare directly | DeepInfra bills most models outside its per-token list by execution time |
| Date-bound prices | A cheap row can double on a known date | Gemini 3.8 Flash is $0.75 / $3.75 through December 31, 2026, then $1.50 / $7.50 |
| Free credit | Rarely changes a production decision | Fireworks gives $1 in free credits |

If two providers serve the same snapshot at the same precision, acceptance rates should match, and the monthly cost decides. If they don't, the cost per accepted output decides, and only your own test set can supply the denominator.

## Where a gateway like laozhang.ai fits

laozhang.ai is the LaoZhang AI Team's own API gateway, so hold it to the same checks as any other provider. It is operated by YingTu Technology Pte. Ltd. in Singapore and states that it is not the official operator of the models it routes to.

What it offers, per its [documentation](https://docs.laozhang.ai/en):

- One key and one base URL (`https://api.laozhang.ai/v1` for OpenAI SDKs) with Anthropic Messages at `/v1/messages` and Gemini-format calls on the same host.
- A prepaid balance. Depending on the model, a request is charged per token or per call, and some token groups apply a price multiplier.
- A Models and pricing page listing model IDs, prices, and groups, plus call logs showing what each request cost.
- Prompt and response content not stored by default. Service metadata such as time, model, usage, status, and latency is kept for billing and troubleshooting.
- Registration requires a Gmail address and is subject to allowlist review.

It fits when you want one balance across OpenAI-, Claude-, and Gemini-format models without separate vendor accounts. It fits less well when you call a single open-weight model at high volume, where a host's per-token row may be simpler, or when you need a contract directly with the model's developer. Before moving traffic, take the model's listed price times your group's multiplier and compare it with the developer's list price. Then confirm the model ID is the version you tested, and check a few calls in the logs. For Claude specifically, see [How to Get Stable Claude Access: Anthropic Direct, Supported Cloud, or the laozhang.ai Gateway?](https://blog.laozhang.ai/en/posts/claude-gateway-laozhang-ai).

## Checklist before moving production traffic

1. **Pin the model.** Record the exact model ID, snapshot date, context length, and precision on each candidate provider.
2. **Copy the live rows.** Take uncached input, cached input, output, and billing unit from each provider's own pricing page on the day you decide.
3. **Apply discounts you can actually use.** Off-peak only counts for traffic that runs in off-peak hours; batch only counts for work that can wait.
4. **Add every fee.** Include credit-purchase fees, BYOK fees above the allowance, group multipliers, and currency conversion.
5. **Price your month.** Use your real input, cache hit rate, output, and retry volumes in the formula above.
6. **Run the same test set.** Send identical tasks to each candidate, count outputs that pass your checks, and compare cost per accepted output.
7. **Check limits and data terms.** Look at rate limits (OpenRouter Standard passes provider limits through), billing for failed calls, region, and retention. For the data side, use [LLM API Data Retention vs Zero Data Retention: A Route-Level Audit](https://blog.laozhang.ai/en/posts/llm-api-data-retention-vs-zero-data-retention).
8. **Ramp behind a cap.** Move a slice of traffic with a spend limit and a fallback. For agents, see [LLM Agent API Spend Kill Switch: Stop Runaway Costs Before the Provider Call](https://blog.laozhang.ai/en/posts/llm-agent-api-spend-kill-switch). Put date-bound prices on your calendar.

## FAQ

### Which LLM API provider is cheapest right now?

None is cheapest across all models. As of September 26, 2026, DeepSeek direct at off-peak rates is the cheapest way to buy DeepSeek V4 Pro and V4.1 Flash among the providers compared here. DeepInfra lists V4 Pro and Kimi K3 below Together, while Together lists Qwen3.7-Max below DeepInfra. For closed models, the developer's own API sets the reference price.

### Is OpenRouter cheaper than calling the provider directly?

Not per token. OpenRouter passes through the upstream provider's price and adds a 5.5% fee on credit purchases (8% on Business), so the same call costs slightly more. It can still cost less overall if one API, automatic fallback, and a single balance save engineering time. On Standard and Business, the first $25,000 of list-price usage each month with your own keys (BYOK) carries no fee.

### How can open-model hosts charge less than the model's developer? Are they quantized?

Hosts run open weights on their own hardware and set their own margins, so their prices differ from the developer's. Some cheap rows are older snapshots, and some listings carry precision labels such as FP8 or FP4. A lower price doesn't prove quantization, and a matching price doesn't prove full precision. Read the model card for snapshot and precision, then run your own test set before switching.

### Are free credits worth choosing a provider for?

Rarely for production. Fireworks offers $1 in free credits, and OpenRouter's Free plan allows 50 requests per day across 25+ free models. That's enough to test, not to run a service. For free tiers with real headroom, compare them in [Best Free AI API Provider: Pick by Limits, Data Use, and Region](https://blog.laozhang.ai/en/posts/free-ai-api-tiers-compared).

### What is the cheapest LLM API for OpenClaw?

OpenClaw takes a base URL and API key, so any provider with an OpenAI-compatible endpoint can be compared the same way. Agent sessions resend long context, so weigh cached-input prices heavily, and set a spend cap before letting an agent loop run. Setup details are in [OpenClaw Complete Guide: API Relay Setup, Model Selection & Cost Optimization (2026)](https://blog.laozhang.ai/en/posts/openclaw-api-guide).
