# Grok API Pricing: Model Rates, Billing Rules, and Monthly Costs

> As of September 26, 2026, grok-4.7 costs $2 input and $6 output per million tokens; grok-4.3 costs $1.25 and $2.50. No free API allowance is listed.

- URL: https://blog.laozhang.ai/en/posts/xai-grok-api-pricing
- Published: 2026-07-02
- Updated: 2026-09-26
- Author: LaoZhang AI Team (https://blog.laozhang.ai/en/about)
- Category: API Pricing
- Tags: Grok API, xAI API, API pricing, grok-4.7, grok-4.3, AI API cost

---
As of September 26, 2026, xAI's flagship API model, `grok-4.7`, costs $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens. The cheaper `grok-4.3` costs $1.25 / $0.20 / $2.50 and has a 1M-token context window. The coding model `grok-build-0.1` costs $1.00 / $0.20 / $2.00. All three prices come from [xAI's pricing page](https://docs.x.ai/developers/pricing), last updated September 21, 2026.

Those numbers hold for requests under 200K tokens on the global endpoint at standard priority. Four things move the bill away from them. Once a prompt reaches 200K tokens, the whole request is billed at double rates. Batch saves 20% only on `grok-4.3` and the `grok-4.20` models, not on `grok-4.7`. Priority Processing doubles every token. The US regional endpoint adds 10%. Server-side tools such as web search are billed per 1,000 calls on top of tokens.

xAI lists no free API allowance, and a SuperGrok or X subscription doesn't pay for API calls. API usage is billed separately through the xAI Console.

## Grok API prices per million tokens

The table below lists every text model on xAI's pricing page. "Short context" applies while the prompt is under 200K tokens; "long context" applies to the whole request once the prompt reaches 200K.

| Model | Context window | Input | Cached input | Output | Long-context input / cached / output |
|---|---:|---:|---:|---:|---:|
| `grok-4.7` | 500K | $2.00 | $0.50 | $6.00 | $4.00 / $1.00 / $12.00 |
| `grok-4.6` | 500K | $2.00 | $0.50 | $6.00 | $4.00 / $1.00 / $12.00 |
| `grok-4.5` | 500K | $2.00 | $0.30 | $6.00 | $4.00 / $0.60 / $12.00 |
| `grok-4.3` | 1M | $1.25 | $0.20 | $2.50 | $2.50 / $0.40 / $5.00 |
| `grok-4.20-0309-reasoning` | 1M | $1.25 | $0.20 | $2.50 | $2.50 / $0.40 / $5.00 |
| `grok-4.20-0309-non-reasoning` | 1M | $1.25 | $0.20 | $2.50 | $2.50 / $0.40 / $5.00 |
| `grok-4.20-multi-agent-0309` | 1M | $1.25 | $0.20 | $2.50 | $2.50 / $0.40 / $5.00 |
| `grok-build-0.1` | 256K | $1.00 | $0.20 | $2.00 | $2.00 / $0.40 / $4.00 |

Prices are in USD per 1M tokens, global endpoint, standard processing, as listed by xAI on September 21, 2026.

Most budgets come down to three rows:

- **`grok-4.7`** is the model xAI describes as its flagship "for code and everything else," with agentic tool calling and configurable reasoning (low, medium, high, xhigh). Pick it when answer quality on hard or tool-heavy tasks matters more than price.
- **`grok-4.3`** costs 37.5% less on input and 58% less on output. It also has the largest window (1M tokens), the cheapest cached input, and the only Batch discount among current models. It's the default for high-volume, cache-heavy, or long-document work.
- **`grok-build-0.1`** is xAI's coding model "for agentic software, engineering, and workflow tasks," and has the lowest rates. Its 256K window is the constraint: a 300K-token repository dump won't fit.

`grok-4.6` costs the same as `grok-4.7`, so it offers no savings. `grok-4.5` has cheaper cached input ($0.30 instead of $0.50) at the same input and output rates. Keep either only if you've pinned it and tested it.

You may also see "Grok 4.7 Fast" at $4 / $1 / $12. It's the same model on faster hardware at twice the price, and it's sold only through Cursor and Grok Build. You can't call it on the public xAI API, so it doesn't belong in an API budget.

## Rules that change the bill

The rates in the table assume a short request, standard priority, no tools, and the global endpoint. Each of these rules can change the total by more than the gap between two models.

| Rule | What xAI charges | Applies to |
|---|---|---|
| Long context | Every token in the request at the long-context rate once the prompt reaches 200K | All text models |
| Batch API | 20% off all token types, results usually within 24 hours | `grok-4.3` and the three `grok-4.20` models only |
| Priority Processing | 2x on every token type, after the cache discount | Chat Completions and Responses API |
| US regional endpoint | 1.1x global rates (`https://us.api.x.ai/v1`) | `grok-4.7` and `grok-4.6` only |
| Web Search / Code Execution | $5 per 1,000 calls | Any model using the tool |
| X Search | $5 per 1,000 posts, $10 per 1,000 profiles returned | Any model using the tool |
| File Attachments search | $10 per 1,000 calls | Any model using the tool |
| Collections Search (RAG) | $2.50 per 1,000 calls | Any model using the tool |
| File / collection storage | $0.025 / $0.10 per GiB per day; downloads $0.20 per GiB | Stored files and collections |
| Usage-guideline violation | $0.05 per request if blocked before generation in the Responses API | Blocked requests |

![One grok-4.7 support-chat month costs $540, rises to $594 on the US endpoint and $1,080 with Priority, has no Batch discount, and drops to $212 on grok-4.3 with Batch; a 300K-token request costs $1.32 at long-context rates instead of $0.66](https://blog.laozhang.ai/posts/en/xai-grok-api-pricing/img/grok-bill-rules.webp)

### Long context bills the whole request, not just the overflow

The 200K line isn't a marginal rate. When a prompt reaches 200K tokens, xAI bills every token in that request at the long-context rate. That includes the first 200K and the output.

One request with 300K input tokens and 10K output tokens shows the effect:

- `grok-4.7` at long-context rates: 0.3 × $4.00 + 0.01 × $12.00 = **$1.32**
- The same request priced with short-context rates, a common spreadsheet mistake: 0.3 × $2.00 + 0.01 × $6.00 = $0.66
- `grok-4.3` at its long-context rates: 0.3 × $2.50 + 0.01 × $5.00 = **$0.80**

If a job regularly passes 200K, trimming it to 190K cuts the price of that request by about half. If it can't be trimmed, `grok-4.3` is the cheaper home, and its 1M window leaves more room than `grok-4.7`'s 500K.

### Batch doesn't discount grok-4.7

A flat 50% Batch discount is a common budgeting assumption, and on xAI it's wrong. The discount varies by model. It is 20% for `grok-4.3` and the `grok-4.20` models, and the pricing page states that "models not listed above have no batch discount." That excludes `grok-4.7`, `grok-4.6`, `grok-4.5`, and `grok-build-0.1`. The `grok-4.7` model page lists Batch API as not supported.

Batch requests usually finish within 24 hours and don't count toward rate limits. A nightly summarization or classification job therefore runs cheapest on `grok-4.3` through Batch. Using `grok-4.7` there costs full price.

### Priority and the US endpoint multiply the rate

Priority Processing doubles input, cached, output, and reasoning tokens. xAI applies the cache discount first and then the 2x multiplier. You pay the priority rate only when the response comes back with `"service_tier": "priority"`. Requests that fall back to the default tier are billed at standard rates.

The US regional endpoint runs inference in the United States and charges 1.1x. For `grok-4.7` that works out to $2.20 / $0.55 / $6.60 under 200K and $4.40 / $1.10 / $13.20 above. It currently serves only `grok-4.7` and `grok-4.6`, so a data-residency requirement also rules out the cheaper models.

### Tools are a separate line item

When a request uses xAI's server-side tools, you pay for the tokens (including search results pulled into the context) plus a per-call fee. X Search is the easy one to underestimate: xAI charges per item returned, and every post in a search or thread fetch counts, including parent and quoted posts. Remote MCP tools have no invocation fee, but the tokens they add are billed.

Image and voice models are priced separately. Image generation runs $0.02 to $0.08 per image depending on model and resolution, video $0.05 to $0.25 per second, and the speech-to-speech voice agent $0.08 per minute.

## Estimate your monthly Grok API bill

Price one month with this formula, then replace each input with numbers from your own logs:

```text
monthly cost =
    new input tokens    / 1,000,000 × input rate
  + cached input tokens / 1,000,000 × cached rate
  + output tokens       / 1,000,000 × output rate      (include reasoning tokens)
  + tool calls          / 1,000     × tool fee
  + storage GiB-days × storage rate
then apply: long-context rates (prompt ≥ 200K), Batch 0.8x (grok-4.3/4.20 only),
            Priority 2x, US endpoint 1.1x
```

The three examples below are hypothetical workloads priced at the September 21, 2026 rates. They exclude retries, which you should add from your own failure rate.

### Support chat: 100,000 conversations a month

Each request carries 1,000 new input tokens, a 2,000-token system prompt served from cache, and 400 output tokens. Over a month that comes to 100M new input tokens, 200M cached input tokens, and 40M output tokens.

| Setup | Calculation | Monthly cost |
|---|---|---:|
| `grok-4.7` | 100 × $2.00 + 200 × $0.50 + 40 × $6.00 | $540 |
| `grok-4.7`, US endpoint | $540 × 1.1 | $594 |
| `grok-4.7`, Priority | $540 × 2 | $1,080 |
| `grok-4.3` | 100 × $1.25 + 200 × $0.20 + 40 × $2.50 | $265 |
| `grok-4.3`, Batch | $265 × 0.8 | $212 |
| `grok-build-0.1` | 100 × $1.00 + 200 × $0.20 + 40 × $2.00 | $220 |

`grok-4.3` costs about half as much as `grok-4.7` here. Part of the gap comes from caching: `grok-4.3`'s cached rate is 84% below its input rate, while `grok-4.7`'s is 75% below. Batch fits only if replies can wait, so it suits ticket triage or after-hours drafts, not live chat.

### Research agent: 1,000 tasks with web search

Each task reads 60K input tokens (search results included), writes 3K output tokens, and makes 5 web searches. That's 60M input, 3M output, and 5,000 search calls a month.

- `grok-4.7`: 60 × $2.00 + 3 × $6.00 = $138 in tokens, plus 5 × $5 = $25 in search fees, for **$163**
- `grok-4.3`: 60 × $1.25 + 3 × $2.50 = $82.50 in tokens, plus the same $25, for **$107.50**

The search fee doesn't shrink with a cheaper model. Capping searches per task is the lever that works on every model.

### Long documents: 300K-token requests

Every request of this size is billed entirely at long-context rates: $1.32 each on `grok-4.7` and $0.80 each on `grok-4.3`. A thousand such requests cost $1,320 or $800 a month. It won't run on `grok-build-0.1`, whose window stops at 256K.

## Is Grok API cheaper than ChatGPT or Claude API?

On list prices, yes for the flagship tier, and by more at the lower tiers. The table prices one month of 100M input and 20M output tokens with no caching, using each vendor's own first-party rate as of September 26, 2026.

| Model | Input / output per 1M | 100M in + 20M out |
|---|---:|---:|
| `grok-build-0.1` | $1.00 / $2.00 | $140 |
| `grok-4.3` | $1.25 / $2.50 | $175 |
| `grok-4.7` | $2.00 / $6.00 | $320 |
| OpenAI `gpt-6-sol` | $2.00 / $10.00 | $400 |
| Claude Sonnet 5 | $2.00 / $10.00 | $400 |
| OpenAI `gpt-6-luna` | $0.10 / $0.50 | $20 |

![Monthly cost of 100M input and 20M output tokens: GPT-6 Sol and Claude Sonnet 5 $400 each, grok-4.7 $320, grok-4.3 $175, grok-build-0.1 $140, GPT-6 Luna $20](https://blog.laozhang.ai/posts/en/xai-grok-api-pricing/img/grok-vs-chatgpt-claude-cost.webp)

`grok-4.7` matches GPT-6 Sol and Claude Sonnet 5 on input and charges 40% less for output, so it gets cheaper the more output a workload produces. `grok-4.3` costs less than half of either. OpenAI's small `gpt-6-luna` is far cheaper than any Grok model, so for simple extraction or classification the cheapest option isn't on xAI.

The comparison has limits. Each vendor uses its own tokenizer, so the same text can produce different token counts. Cache and Batch rules also differ by vendor, and no Batch discount applies to `grok-4.7` at all. For a wider comparison by input and output mix, see [AI API Price Comparison: Cheapest Model by Input and Output Mix](https://blog.laozhang.ai/en/posts/cheapest-llm-models). The OpenAI and Anthropic tiers are covered in [Claude API vs OpenAI API Pricing: When Each One Costs Less](https://blog.laozhang.ai/en/posts/claude-api-vs-openai-api-pricing).

## Is the Grok API free?

xAI's pricing, models, and rate-limit pages list no free API credit or free token allowance. The "Tier 0" on the rate-limit page is the $0 starting point for rate limits, not free usage. Promotions may appear in individual consoles, but none is part of the published price list.

Consumer plans are a different product. Free, SuperGrok, Business, and Enterprise on x.ai cover the Grok app. API calls are billed per token to your xAI Console team, whatever app plan you hold. "SuperGrok API" isn't a product.

Two more sources of confusion:

- **Resellers and gateways.** OpenRouter and similar gateways, along with regional resellers, sell Grok access at their own prices and terms. Their rows may match xAI's or differ from them. Either way, they're a separate agreement with that company, not xAI's price. [Cheapest LLM API Provider: Compare Price, Quality, Latency, and Gateway Risk in 2026](https://blog.laozhang.ai/en/posts/cheapest-llm-api-provider) covers what to check before routing through one.
- **Groq is not Grok.** Groq (groq.com) is an unrelated inference company with its own pricing and free tier.

If you need a genuinely free API to prototype with, [Best Free AI API Provider: Pick by Limits, Data Use, and Region](https://blog.laozhang.ai/en/posts/free-ai-api-tiers-compared) compares the options that do publish one.

## Keep Grok API spend under control

**Read the cost from every response.** xAI returns the billed cost of each request in `usage.cost_in_usd_ticks`, after cache discounts and including server-side tool fees. One US dollar equals 10,000,000,000 ticks. Log it next to the model ID and task type, and your worksheet can be checked against real spend from the first day:

```python
import os
from openai import OpenAI

client = OpenAI(api_key=os.getenv("XAI_API_KEY"), base_url="https://api.x.ai/v1")

resp = client.chat.completions.create(
    model="grok-4.3",
    messages=[{"role": "user", "content": "Summarize this ticket in two lines."}],
)
cost_usd = resp.usage.cost_in_usd_ticks / 1e10
print(f"{resp.usage.prompt_tokens} in, {resp.usage.completion_tokens} out, ${cost_usd:.6f}")
```

For streaming with the OpenAI SDK, set `stream_options={"include_usage": True}`, and the cost arrives in the final chunk. The value is per request. Add turns up yourself for a conversation or an agent loop. See [xAI's cost tracking guide](https://docs.x.ai/developers/cost-tracking) for the xAI SDK version.

**Know how rate limits grow with spend.** Tiers are set by cumulative API spend since January 1, 2026: Tier 1 at $50, Tier 2 at $250, Tier 3 at $1,000, and Tier 4 at $5,000. Prepaid credit purchases and paid invoices both count, and tiers never drop. For `grok-4.7`, tokens per minute rise from 50M at Tier 0 to 100M at Tier 4. Cached tokens are cheaper, but they still count toward that per-minute limit. Going over returns HTTP 429, so add exponential backoff before a traffic spike turns into failed tasks. Current tier tables are on the [xAI rate limits page](https://docs.x.ai/developers/rate-limits).

**Put a stop rule in code.** Before scaling, run a few hundred real requests and compare the logged cost per task with your estimate. Then set a threshold, for example 30% over the worksheet on cost per task, tool calls per task, or share of prompts over 200K, and pause the job when it trips. A runaway agent loop costs the most on `grok-4.7` with Priority, where each token is billed at twice the flagship rate.

For model IDs, aliases, and a migration test plan on the cheaper model, see the [Grok 4.3 API Guide: Model ID, Pricing, Migration, and Test Plan](https://blog.laozhang.ai/en/posts/grok-4-3).

## FAQ

### How much does the Grok API cost per million tokens?

`grok-4.7` costs $2.00 input, $0.50 cached input, and $6.00 output per million tokens. `grok-4.3` costs $1.25 / $0.20 / $2.50, and `grok-build-0.1` costs $1.00 / $0.20 / $2.00. Requests whose prompt reaches 200K tokens are billed entirely at double those rates. These are xAI's list prices as of September 21, 2026.

### Is the Grok API cheaper than the ChatGPT API?

For comparable flagships, yes on output. `grok-4.7` and OpenAI's `gpt-6-sol` both charge $2 per million input tokens, but Grok charges $6 per million output tokens against Sol's $10. On 100M input and 20M output tokens, that's $320 against $400. OpenAI's `gpt-6-luna` is much cheaper than any Grok model for light tasks.

### Does a SuperGrok subscription include API access?

No. SuperGrok and the other x.ai plans cover the Grok app. API usage is billed per token to your xAI Console team, and xAI's pricing pages don't list API credits as part of any plan.

### Does the Batch API make grok-4.7 cheaper?

No. xAI's 20% Batch discount applies only to `grok-4.3` and the three `grok-4.20` models, and the `grok-4.7` model page lists Batch as not supported. For batch jobs, `grok-4.3` at 80% of its rate is the cheapest current Grok option.

### Which Grok model is cheapest for coding?

`grok-build-0.1`, at $1.00 input and $2.00 output per million tokens, is both the cheapest row and the model xAI built for agentic coding. Its 256K context window is the catch. For larger codebases in one request, `grok-4.3` holds up to 1M tokens at $1.25 / $2.50.
