Skip to main content

Cheapest LLM API Provider: Where the Same Model Costs Less

No provider wins for every model. As of September 26, 2026, DeepSeek direct off-peak costs half of hosts charging its peak rate; check versions and router fees.

LaoZhang AI TeamPublishedUpdated 13 min read
On this page
Monthly cost of the same model at different providers: Kimi K3 cheapest on DeepInfra, Qwen3.7-Max on Together, DeepSeek V4 Pro direct off-peak

No single company sells LLM inference cheapest across the board. The cheapest provider depends on which model you're buying, what hours your traffic runs, and whether two price rows are even the same model version. Prices below are as of September 26, 2026, in USD per 1 million tokens.

Three patterns decide most cases:

  • Closed models (GPT-6, Claude, Gemini, Grok) are priced by their developer. A router passes that price through and adds its own fee. A reseller sets its own price, which you check model by model.
  • Open-weight models (DeepSeek, Kimi, Qwen, GLM, gpt-oss) are sold by many hosts, and the winner flips by model. DeepInfra lists DeepSeek V4 Pro and Kimi K3 below Together; Together lists Qwen3.7-Max below DeepInfra.
  • DeepSeek's own API charges half price outside two weekday UTC windows. Together's DeepSeek rows match DeepSeek's peak price at all hours, so off-peak traffic is cheapest when you buy direct.
Your situationUsually cheapestCheck before switching
You use a closed modelThe developer's own API, with Batch or Flex for work that can waitWhether a router fee or reseller multiplier buys you something you need
You use DeepSeek and most calls land outside 01:00–04:00 and 06:00–10:00 UTC on weekdaysDeepSeek direct at off-peak ratesYour real hour-by-hour traffic mix
You need DeepSeek V4 Pro during peak hoursDeepInfra ($1.30 input / $2.60 output) over Together or DeepSeek peak ($1.32 / $3.96)That DeepInfra's unlabeled V4 Pro snapshot matches the one you tested
You use another open-weight modelWhichever host lists that exact model lower; it variesSnapshot date, context length, and precision label
You switch models often or want automatic fallbackA router such as OpenRouterA 5.5% credit-purchase fee on the Standard plan

Which model to buy is a separate decision. AI API Price Comparison: Cheapest Model by Input and Output Mix covers that. Everything below assumes you already know the model and want the cheapest place to call it.

Four kinds of LLM API providers

"Provider" covers four different businesses. They sell different things, so their prices mean different things.

TypeExamplesWhat you're paying forWhere the price comes fromWhat to watch
First-party APIOpenAI, Anthropic, Google, xAI, DeepSeekThe model from the company that trains it, including its batch, flex, or off-peak discountsThe developer's pricing pageDate-bound prices and retired model names
Open-model hostTogether, DeepInfra, Fireworks, Groq, SiliconFlowOpen weights served on the host's hardwareThe host's own table, set model by modelSnapshot version, precision, and billing unit
RouterOpenRouterOne API and one balance across many upstream providersThe upstream provider's price, plus a platform feeCredit-purchase fees, BYOK fees, plan limits
Gateway or resellerlaozhang.ai, regional resellersOne compatible endpoint, prepaid balance, often local paymentThe gateway's own table, sometimes with group multipliersWhich model actually answers, and who supports you when upstream fails

The first-party row is the reference for closed models. As of September 26, 2026, OpenAI lists GPT-6 Luna at $0.10 input and $0.50 output, with Batch and Flex at half price. Google lists Gemini 3.1 Flash-Lite at $0.25 / $1.50. Anthropic lists Claude Haiku 4.5 at $1 / $5, and xAI lists grok-4.7 at $2 / $6. Open-model hosts don't serve these models. Other sellers, such as cloud platforms, routers, and resellers, either pass the developer's price through or set their own, which you verify on their page. For side-by-side closed-model pricing, see Claude API vs OpenAI API Pricing: When Each One Costs Less.

Open-weight models are where shopping across providers pays. Each host picks which snapshot it runs, at what precision, and at what margin. Fireworks bills serverless models per token, postpaid, with separate Standard, Priority, and Fast tiers, so even one host can list several prices for one model. DeepInfra says some of its language models are priced per token and most of its other models are billed by execution time.

Same model, different monthly bill

The table prices one month of 100 million input tokens and 20 million output tokens with no caching. Each figure is input millions × input rate + output millions × output rate, so you can redo it with your own volumes.

ModelProvider and rowInput / output per 1MMonth (100M in, 20M out)
DeepSeek V4 ProDeepSeek direct, off-peak$0.66 / $1.98$105.60
DeepSeek V4 ProDeepSeek direct, peak$1.32 / $3.96$211.20
DeepSeek V4 ProTogether (V4 Pro 0813)$1.32 / $3.96$211.20
DeepSeek V4 ProDeepInfra (DeepSeek-V4-Pro, no date label)$1.30 / $2.60$182.00
DeepSeek V4.1 FlashDeepSeek direct, off-peak$0.15 / $0.60$27.00
DeepSeek V4.1 FlashDeepSeek direct, peak$0.30 / $1.20$54.00
DeepSeek V4.1 FlashTogether$0.30 / $1.20$54.00
Kimi K3Together$3.00 / $15.00$600.00
Kimi K3DeepInfra$2.85 / $14.25$570.00
Qwen3.7-MaxTogether$1.50 / $4.50$240.00
Qwen3.7-MaxDeepInfra$2.50 / $7.50$400.00

Prices from DeepSeek pricing, Together pricing, and DeepInfra pricing.

Three conclusions hold for this mix:

  • Neither host is cheaper overall. DeepInfra saves $30 a month on Kimi K3 and $29.20 on V4 Pro. Together saves $160 on Qwen3.7-Max.
  • Time of day beats host choice for DeepSeek. Off-peak direct V4 Pro costs $105.60, well under DeepInfra's $182.00. During peak hours, DeepInfra is lower than both DeepSeek and Together, provided its snapshot passes your tests.
  • Caching widens the gap. With 60% of input served from cache, a V4.1 Flash month costs $18.18 at DeepSeek off-peak and $36.36 on Together: 40M × $0.15 + 60M × $0.003 + 20M × $0.60, against 40M × $0.30 + 60M × $0.006 + 20M × $1.20.

Bar chart of monthly DeepSeek bills for 100M input and 20M output tokens: V4 Pro costs $105.60 direct off-peak, $182.00 on DeepInfra, and $211.20 at DeepSeek peak or on Together

DeepSeek's peak windows are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays. During US daylight saving time that is 9 PM–midnight and 2–6 AM Eastern. A 9-to-5 workday on either US coast falls entirely outside the peak windows. For a job-by-job scheduling test, see DeepSeek V4 Peak Pricing: Calculate the Real Cost Before You Reschedule.

Four ways a cheap price row misleads

Four traps behind a cheap LLM price row: a host charging the peak rate at all hours, an older model snapshot, a router credit fee, and a ranking written by a listed provider

The host charges the developer's peak price

Together's DeepSeek V4 Pro 0813 row ($1.32 / $3.96) and V4.1 Flash row ($0.30 / $1.20) are identical to DeepSeek's peak prices. Together's page shows no off-peak discount. A host that matches the peak price looks competitive in a table that only shows DeepSeek's peak row, but it costs double whenever DeepSeek's off-peak price applies. A host can still be worth it for reasons other than price, such as one account for several open models. Just price it against the rate you would actually pay direct.

The cheaper row is an older snapshot

DeepSeek retired the deepseek-v4-flash API name. Calls to it are now served by V4.1-Flash at the Flash price. Hosts still sell the earlier open-weight snapshot: Together lists DeepSeek V4 Flash 0731 at $0.14 / $0.28, and DeepInfra lists DeepSeek-V4-Flash-0731 at $0.06 / $0.18. DeepInfra also has a row simply labeled DeepSeek-V4-Flash at $0.09 / $0.18, and its page doesn't say which snapshot that is.

Those rows are real and cheap. On the month above, the 0731 snapshot costs $19.60 on Together and $9.60 on DeepInfra, versus $27.00 for V4.1 Flash off-peak direct. They are not the model DeepSeek serves today, though. An older snapshot can still win on your workload. Suppose, hypothetically, that 30% of its outputs fail your checks and you rerun those tasks on V4.1 Flash off-peak. The month then costs about $9.60 + 0.30 × $27.00 = $17.70, still below $27.00. The mistake is assuming the rows are interchangeable, not choosing the older snapshot after testing it.

Precision matters in the same way. Some host listings carry labels such as FP8 or FP4 in the model name. Read the model card for the snapshot date, context length, and precision before comparing prices.

The router adds a fee on top of pass-through prices

OpenRouter says it passes through the underlying provider's price without markup and charges a fee when you buy credits: 5.5% on the Standard plan and 8% on Business, with discounts negotiated for Enterprise. A $500 top-up on Standard carries a $27.50 fee. The V4 Pro month that costs $211.20 at a provider's list price therefore costs about $222.82 through OpenRouter credits.

With your own provider keys (BYOK), Standard and Business include $25,000 of list-price usage per month with no fee, then charge 5%. The Free plan has no platform fee but limits you to 25+ free models and 50 requests per day, with no BYOK. The pass-through price is whichever upstream provider serves the call, so check which provider that is on the model page rather than assuming the lowest row in this table. See OpenRouter pricing and the OpenRouter FAQ.

A router pays for itself when switching models, falling back across providers, or consolidating billing saves more engineering time than the fee costs. It is not a way to get a lower per-token price than the provider it routes to.

The ranking comes from a provider on the list

"Cheapest provider" lists written by a host that includes itself are marketing. So are router pages quoting a headline price per million tokens without naming the model and version. A useful comparison prices the same model snapshot on each provider, shows input, cached input, and output separately, and dates the prices. If a ranking doesn't do that, use it only as a list of names to check.

Effective cost: the calculation that decides it

A price table is an input, not the answer. Price a real month of your own traffic, then divide by the work that actually passed.

monthly cost = ( uncached input M × input rate
               + cached input M  × cached rate
               + output M        × output rate ) × (1 + platform fee)

cost per accepted output = monthly cost ÷ outputs that pass your checks

Count retries in the token volumes, because every failed call you retry is billed again. Then adjust for the factors that change which provider wins:

FactorWhy it moves the answerExample as of September 26, 2026
Time or mode discountsThe same model can cost half as much at a different hour or in batchDeepSeek off-peak is half of peak; OpenAI Batch and Flex and Gemini Batch are half price
Cached inputLong, repeated context can dominate agent and RAG billsV4.1 Flash cached input is $0.003 at DeepSeek off-peak and $0.006 on Together
Platform feesA fee on credits applies to every token you buyOpenRouter: 5.5% Standard, 8% Business
Version and precisionA cheaper row may be a different modelV4 Flash 0731 rows versus DeepSeek's current V4.1-Flash
Billing unitPer token, per call, and per second of execution don't compare directlyDeepInfra bills most models outside its per-token list by execution time
Date-bound pricesA cheap row can double on a known dateGemini 3.8 Flash is $0.75 / $3.75 through December 31, 2026, then $1.50 / $7.50
Free creditRarely changes a production decisionFireworks gives $1 in free credits

If two providers serve the same snapshot at the same precision, acceptance rates should match, and the monthly cost decides. If they don't, the cost per accepted output decides, and only your own test set can supply the denominator.

Where a gateway like laozhang.ai fits

laozhang.ai is the LaoZhang AI Team's own API gateway, so hold it to the same checks as any other provider. It is operated by YingTu Technology Pte. Ltd. in Singapore and states that it is not the official operator of the models it routes to.

What it offers, per its documentation:

  • One key and one base URL (https://api.laozhang.ai/v1 for OpenAI SDKs) with Anthropic Messages at /v1/messages and Gemini-format calls on the same host.
  • A prepaid balance. Depending on the model, a request is charged per token or per call, and some token groups apply a price multiplier.
  • A Models and pricing page listing model IDs, prices, and groups, plus call logs showing what each request cost.
  • Prompt and response content not stored by default. Service metadata such as time, model, usage, status, and latency is kept for billing and troubleshooting.
  • Registration requires a Gmail address and is subject to allowlist review.

It fits when you want one balance across OpenAI-, Claude-, and Gemini-format models without separate vendor accounts. It fits less well when you call a single open-weight model at high volume, where a host's per-token row may be simpler, or when you need a contract directly with the model's developer. Before moving traffic, take the model's listed price times your group's multiplier and compare it with the developer's list price. Then confirm the model ID is the version you tested, and check a few calls in the logs. For Claude specifically, see How to Get Stable Claude Access: Anthropic Direct, Supported Cloud, or the laozhang.ai Gateway?.

Checklist before moving production traffic

  1. Pin the model. Record the exact model ID, snapshot date, context length, and precision on each candidate provider.
  2. Copy the live rows. Take uncached input, cached input, output, and billing unit from each provider's own pricing page on the day you decide.
  3. Apply discounts you can actually use. Off-peak only counts for traffic that runs in off-peak hours; batch only counts for work that can wait.
  4. Add every fee. Include credit-purchase fees, BYOK fees above the allowance, group multipliers, and currency conversion.
  5. Price your month. Use your real input, cache hit rate, output, and retry volumes in the formula above.
  6. Run the same test set. Send identical tasks to each candidate, count outputs that pass your checks, and compare cost per accepted output.
  7. Check limits and data terms. Look at rate limits (OpenRouter Standard passes provider limits through), billing for failed calls, region, and retention. For the data side, use LLM API Data Retention vs Zero Data Retention: A Route-Level Audit.
  8. Ramp behind a cap. Move a slice of traffic with a spend limit and a fallback. For agents, see LLM Agent API Spend Kill Switch: Stop Runaway Costs Before the Provider Call. Put date-bound prices on your calendar.

FAQ

Which LLM API provider is cheapest right now?

None is cheapest across all models. As of September 26, 2026, DeepSeek direct at off-peak rates is the cheapest way to buy DeepSeek V4 Pro and V4.1 Flash among the providers compared here. DeepInfra lists V4 Pro and Kimi K3 below Together, while Together lists Qwen3.7-Max below DeepInfra. For closed models, the developer's own API sets the reference price.

Is OpenRouter cheaper than calling the provider directly?

Not per token. OpenRouter passes through the upstream provider's price and adds a 5.5% fee on credit purchases (8% on Business), so the same call costs slightly more. It can still cost less overall if one API, automatic fallback, and a single balance save engineering time. On Standard and Business, the first $25,000 of list-price usage each month with your own keys (BYOK) carries no fee.

How can open-model hosts charge less than the model's developer? Are they quantized?

Hosts run open weights on their own hardware and set their own margins, so their prices differ from the developer's. Some cheap rows are older snapshots, and some listings carry precision labels such as FP8 or FP4. A lower price doesn't prove quantization, and a matching price doesn't prove full precision. Read the model card for snapshot and precision, then run your own test set before switching.

Are free credits worth choosing a provider for?

Rarely for production. Fireworks offers $1 in free credits, and OpenRouter's Free plan allows 50 requests per day across 25+ free models. That's enough to test, not to run a service. For free tiers with real headroom, compare them in Best Free AI API Provider: Pick by Limits, Data Use, and Region.

What is the cheapest LLM API for OpenClaw?

OpenClaw takes a base URL and API key, so any provider with an OpenAI-compatible endpoint can be compared the same way. Agent sessions resend long context, so weigh cached-input prices heavily, and set a spend cap before letting an agent loop run. Setup details are in OpenClaw Complete Guide: API Relay Setup, Model Selection & Cost Optimization (2026).