Skip to main content

Grok API Pricing: Model Rates, Billing Rules, and Monthly Costs

As of September 26, 2026, grok-4.7 costs $2 input and $6 output per million tokens; grok-4.3 costs $1.25 and $2.50. No free API allowance is listed.

LaoZhang AI TeamPublishedUpdated 12 min read
On this page
Grok API pricing cover: grok-4.7 at $2.00 input and $6.00 output per million tokens, grok-4.3 at $1.25 and $2.50, grok-build-0.1 at $1.00 and $2.00, with prompts of 200K or more billed at 2x and no free API allowance

As of September 26, 2026, xAI's flagship API model, grok-4.7, costs $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens. The cheaper grok-4.3 costs $1.25 / $0.20 / $2.50 and has a 1M-token context window. The coding model grok-build-0.1 costs $1.00 / $0.20 / $2.00. All three prices come from xAI's pricing page, last updated September 21, 2026.

Those numbers hold for requests under 200K tokens on the global endpoint at standard priority. Four things move the bill away from them. Once a prompt reaches 200K tokens, the whole request is billed at double rates. Batch saves 20% only on grok-4.3 and the grok-4.20 models, not on grok-4.7. Priority Processing doubles every token. The US regional endpoint adds 10%. Server-side tools such as web search are billed per 1,000 calls on top of tokens.

xAI lists no free API allowance, and a SuperGrok or X subscription doesn't pay for API calls. API usage is billed separately through the xAI Console.

Grok API prices per million tokens

The table below lists every text model on xAI's pricing page. "Short context" applies while the prompt is under 200K tokens; "long context" applies to the whole request once the prompt reaches 200K.

ModelContext windowInputCached inputOutputLong-context input / cached / output
grok-4.7500K$2.00$0.50$6.00$4.00 / $1.00 / $12.00
grok-4.6500K$2.00$0.50$6.00$4.00 / $1.00 / $12.00
grok-4.5500K$2.00$0.30$6.00$4.00 / $0.60 / $12.00
grok-4.31M$1.25$0.20$2.50$2.50 / $0.40 / $5.00
grok-4.20-0309-reasoning1M$1.25$0.20$2.50$2.50 / $0.40 / $5.00
grok-4.20-0309-non-reasoning1M$1.25$0.20$2.50$2.50 / $0.40 / $5.00
grok-4.20-multi-agent-03091M$1.25$0.20$2.50$2.50 / $0.40 / $5.00
grok-build-0.1256K$1.00$0.20$2.00$2.00 / $0.40 / $4.00

Prices are in USD per 1M tokens, global endpoint, standard processing, as listed by xAI on September 21, 2026.

Most budgets come down to three rows:

  • grok-4.7 is the model xAI describes as its flagship "for code and everything else," with agentic tool calling and configurable reasoning (low, medium, high, xhigh). Pick it when answer quality on hard or tool-heavy tasks matters more than price.
  • grok-4.3 costs 37.5% less on input and 58% less on output. It also has the largest window (1M tokens), the cheapest cached input, and the only Batch discount among current models. It's the default for high-volume, cache-heavy, or long-document work.
  • grok-build-0.1 is xAI's coding model "for agentic software, engineering, and workflow tasks," and has the lowest rates. Its 256K window is the constraint: a 300K-token repository dump won't fit.

grok-4.6 costs the same as grok-4.7, so it offers no savings. grok-4.5 has cheaper cached input ($0.30 instead of $0.50) at the same input and output rates. Keep either only if you've pinned it and tested it.

You may also see "Grok 4.7 Fast" at $4 / $1 / $12. It's the same model on faster hardware at twice the price, and it's sold only through Cursor and Grok Build. You can't call it on the public xAI API, so it doesn't belong in an API budget.

Rules that change the bill

The rates in the table assume a short request, standard priority, no tools, and the global endpoint. Each of these rules can change the total by more than the gap between two models.

RuleWhat xAI chargesApplies to
Long contextEvery token in the request at the long-context rate once the prompt reaches 200KAll text models
Batch API20% off all token types, results usually within 24 hoursgrok-4.3 and the three grok-4.20 models only
Priority Processing2x on every token type, after the cache discountChat Completions and Responses API
US regional endpoint1.1x global rates (https://us.api.x.ai/v1)grok-4.7 and grok-4.6 only
Web Search / Code Execution$5 per 1,000 callsAny model using the tool
X Search$5 per 1,000 posts, $10 per 1,000 profiles returnedAny model using the tool
File Attachments search$10 per 1,000 callsAny model using the tool
Collections Search (RAG)$2.50 per 1,000 callsAny model using the tool
File / collection storage$0.025 / $0.10 per GiB per day; downloads $0.20 per GiBStored files and collections
Usage-guideline violation$0.05 per request if blocked before generation in the Responses APIBlocked requests

One grok-4.7 support-chat month costs $540, rises to $594 on the US endpoint and $1,080 with Priority, has no Batch discount, and drops to $212 on grok-4.3 with Batch; a 300K-token request costs $1.32 at long-context rates instead of $0.66

Long context bills the whole request, not just the overflow

The 200K line isn't a marginal rate. When a prompt reaches 200K tokens, xAI bills every token in that request at the long-context rate. That includes the first 200K and the output.

One request with 300K input tokens and 10K output tokens shows the effect:

  • grok-4.7 at long-context rates: 0.3 × $4.00 + 0.01 × $12.00 = $1.32
  • The same request priced with short-context rates, a common spreadsheet mistake: 0.3 × $2.00 + 0.01 × $6.00 = $0.66
  • grok-4.3 at its long-context rates: 0.3 × $2.50 + 0.01 × $5.00 = $0.80

If a job regularly passes 200K, trimming it to 190K cuts the price of that request by about half. If it can't be trimmed, grok-4.3 is the cheaper home, and its 1M window leaves more room than grok-4.7's 500K.

Batch doesn't discount grok-4.7

A flat 50% Batch discount is a common budgeting assumption, and on xAI it's wrong. The discount varies by model. It is 20% for grok-4.3 and the grok-4.20 models, and the pricing page states that "models not listed above have no batch discount." That excludes grok-4.7, grok-4.6, grok-4.5, and grok-build-0.1. The grok-4.7 model page lists Batch API as not supported.

Batch requests usually finish within 24 hours and don't count toward rate limits. A nightly summarization or classification job therefore runs cheapest on grok-4.3 through Batch. Using grok-4.7 there costs full price.

Priority and the US endpoint multiply the rate

Priority Processing doubles input, cached, output, and reasoning tokens. xAI applies the cache discount first and then the 2x multiplier. You pay the priority rate only when the response comes back with "service_tier": "priority". Requests that fall back to the default tier are billed at standard rates.

The US regional endpoint runs inference in the United States and charges 1.1x. For grok-4.7 that works out to $2.20 / $0.55 / $6.60 under 200K and $4.40 / $1.10 / $13.20 above. It currently serves only grok-4.7 and grok-4.6, so a data-residency requirement also rules out the cheaper models.

Tools are a separate line item

When a request uses xAI's server-side tools, you pay for the tokens (including search results pulled into the context) plus a per-call fee. X Search is the easy one to underestimate: xAI charges per item returned, and every post in a search or thread fetch counts, including parent and quoted posts. Remote MCP tools have no invocation fee, but the tokens they add are billed.

Image and voice models are priced separately. Image generation runs $0.02 to $0.08 per image depending on model and resolution, video $0.05 to $0.25 per second, and the speech-to-speech voice agent $0.08 per minute.

Estimate your monthly Grok API bill

Price one month with this formula, then replace each input with numbers from your own logs:

monthly cost =
    new input tokens    / 1,000,000 × input rate
  + cached input tokens / 1,000,000 × cached rate
  + output tokens       / 1,000,000 × output rate      (include reasoning tokens)
  + tool calls          / 1,000     × tool fee
  + storage GiB-days × storage rate
then apply: long-context rates (prompt ≥ 200K), Batch 0.8x (grok-4.3/4.20 only),
            Priority 2x, US endpoint 1.1x

The three examples below are hypothetical workloads priced at the September 21, 2026 rates. They exclude retries, which you should add from your own failure rate.

Support chat: 100,000 conversations a month

Each request carries 1,000 new input tokens, a 2,000-token system prompt served from cache, and 400 output tokens. Over a month that comes to 100M new input tokens, 200M cached input tokens, and 40M output tokens.

SetupCalculationMonthly cost
grok-4.7100 × $2.00 + 200 × $0.50 + 40 × $6.00$540
grok-4.7, US endpoint$540 × 1.1$594
grok-4.7, Priority$540 × 2$1,080
grok-4.3100 × $1.25 + 200 × $0.20 + 40 × $2.50$265
grok-4.3, Batch$265 × 0.8$212
grok-build-0.1100 × $1.00 + 200 × $0.20 + 40 × $2.00$220

grok-4.3 costs about half as much as grok-4.7 here. Part of the gap comes from caching: grok-4.3's cached rate is 84% below its input rate, while grok-4.7's is 75% below. Batch fits only if replies can wait, so it suits ticket triage or after-hours drafts, not live chat.

Each task reads 60K input tokens (search results included), writes 3K output tokens, and makes 5 web searches. That's 60M input, 3M output, and 5,000 search calls a month.

  • grok-4.7: 60 × $2.00 + 3 × $6.00 = $138 in tokens, plus 5 × $5 = $25 in search fees, for $163
  • grok-4.3: 60 × $1.25 + 3 × $2.50 = $82.50 in tokens, plus the same $25, for $107.50

The search fee doesn't shrink with a cheaper model. Capping searches per task is the lever that works on every model.

Long documents: 300K-token requests

Every request of this size is billed entirely at long-context rates: $1.32 each on grok-4.7 and $0.80 each on grok-4.3. A thousand such requests cost $1,320 or $800 a month. It won't run on grok-build-0.1, whose window stops at 256K.

Is Grok API cheaper than ChatGPT or Claude API?

On list prices, yes for the flagship tier, and by more at the lower tiers. The table prices one month of 100M input and 20M output tokens with no caching, using each vendor's own first-party rate as of September 26, 2026.

ModelInput / output per 1M100M in + 20M out
grok-build-0.1$1.00 / $2.00$140
grok-4.3$1.25 / $2.50$175
grok-4.7$2.00 / $6.00$320
OpenAI gpt-6-sol$2.00 / $10.00$400
Claude Sonnet 5$2.00 / $10.00$400
OpenAI gpt-6-luna$0.10 / $0.50$20

Monthly cost of 100M input and 20M output tokens: GPT-6 Sol and Claude Sonnet 5 $400 each, grok-4.7 $320, grok-4.3 $175, grok-build-0.1 $140, GPT-6 Luna $20

grok-4.7 matches GPT-6 Sol and Claude Sonnet 5 on input and charges 40% less for output, so it gets cheaper the more output a workload produces. grok-4.3 costs less than half of either. OpenAI's small gpt-6-luna is far cheaper than any Grok model, so for simple extraction or classification the cheapest option isn't on xAI.

The comparison has limits. Each vendor uses its own tokenizer, so the same text can produce different token counts. Cache and Batch rules also differ by vendor, and no Batch discount applies to grok-4.7 at all. For a wider comparison by input and output mix, see AI API Price Comparison: Cheapest Model by Input and Output Mix. The OpenAI and Anthropic tiers are covered in Claude API vs OpenAI API Pricing: When Each One Costs Less.

Is the Grok API free?

xAI's pricing, models, and rate-limit pages list no free API credit or free token allowance. The "Tier 0" on the rate-limit page is the $0 starting point for rate limits, not free usage. Promotions may appear in individual consoles, but none is part of the published price list.

Consumer plans are a different product. Free, SuperGrok, Business, and Enterprise on x.ai cover the Grok app. API calls are billed per token to your xAI Console team, whatever app plan you hold. "SuperGrok API" isn't a product.

Two more sources of confusion:

  • Resellers and gateways. OpenRouter and similar gateways, along with regional resellers, sell Grok access at their own prices and terms. Their rows may match xAI's or differ from them. Either way, they're a separate agreement with that company, not xAI's price. Cheapest LLM API Provider: Compare Price, Quality, Latency, and Gateway Risk in 2026 covers what to check before routing through one.
  • Groq is not Grok. Groq (groq.com) is an unrelated inference company with its own pricing and free tier.

If you need a genuinely free API to prototype with, Best Free AI API Provider: Pick by Limits, Data Use, and Region compares the options that do publish one.

Keep Grok API spend under control

Read the cost from every response. xAI returns the billed cost of each request in usage.cost_in_usd_ticks, after cache discounts and including server-side tool fees. One US dollar equals 10,000,000,000 ticks. Log it next to the model ID and task type, and your worksheet can be checked against real spend from the first day:

python
import os
from openai import OpenAI

client = OpenAI(api_key=os.getenv("XAI_API_KEY"), base_url="https://api.x.ai/v1")

resp = client.chat.completions.create(
    model="grok-4.3",
    messages=[{"role": "user", "content": "Summarize this ticket in two lines."}],
)
cost_usd = resp.usage.cost_in_usd_ticks / 1e10
print(f"{resp.usage.prompt_tokens} in, {resp.usage.completion_tokens} out, ${cost_usd:.6f}")

For streaming with the OpenAI SDK, set stream_options={"include_usage": True}, and the cost arrives in the final chunk. The value is per request. Add turns up yourself for a conversation or an agent loop. See xAI's cost tracking guide for the xAI SDK version.

Know how rate limits grow with spend. Tiers are set by cumulative API spend since January 1, 2026: Tier 1 at $50, Tier 2 at $250, Tier 3 at $1,000, and Tier 4 at $5,000. Prepaid credit purchases and paid invoices both count, and tiers never drop. For grok-4.7, tokens per minute rise from 50M at Tier 0 to 100M at Tier 4. Cached tokens are cheaper, but they still count toward that per-minute limit. Going over returns HTTP 429, so add exponential backoff before a traffic spike turns into failed tasks. Current tier tables are on the xAI rate limits page.

Put a stop rule in code. Before scaling, run a few hundred real requests and compare the logged cost per task with your estimate. Then set a threshold, for example 30% over the worksheet on cost per task, tool calls per task, or share of prompts over 200K, and pause the job when it trips. A runaway agent loop costs the most on grok-4.7 with Priority, where each token is billed at twice the flagship rate.

For model IDs, aliases, and a migration test plan on the cheaper model, see the Grok 4.3 API Guide: Model ID, Pricing, Migration, and Test Plan.

FAQ

How much does the Grok API cost per million tokens?

grok-4.7 costs $2.00 input, $0.50 cached input, and $6.00 output per million tokens. grok-4.3 costs $1.25 / $0.20 / $2.50, and grok-build-0.1 costs $1.00 / $0.20 / $2.00. Requests whose prompt reaches 200K tokens are billed entirely at double those rates. These are xAI's list prices as of September 21, 2026.

Is the Grok API cheaper than the ChatGPT API?

For comparable flagships, yes on output. grok-4.7 and OpenAI's gpt-6-sol both charge $2 per million input tokens, but Grok charges $6 per million output tokens against Sol's $10. On 100M input and 20M output tokens, that's $320 against $400. OpenAI's gpt-6-luna is much cheaper than any Grok model for light tasks.

Does a SuperGrok subscription include API access?

No. SuperGrok and the other x.ai plans cover the Grok app. API usage is billed per token to your xAI Console team, and xAI's pricing pages don't list API credits as part of any plan.

Does the Batch API make grok-4.7 cheaper?

No. xAI's 20% Batch discount applies only to grok-4.3 and the three grok-4.20 models, and the grok-4.7 model page lists Batch as not supported. For batch jobs, grok-4.3 at 80% of its rate is the cheapest current Grok option.

Which Grok model is cheapest for coding?

grok-build-0.1, at $1.00 input and $2.00 output per million tokens, is both the cheapest row and the model xAI built for agentic coding. Its 256K context window is the catch. For larger codebases in one request, grok-4.3 holds up to 1M tokens at $1.25 / $2.50.