# Best Free AI API Provider: Pick by Limits, Data Use, and Region

> Start with Gemini's free tier unless your prompts are private or you're outside its regions; add Groq or OpenRouter for headroom. As of September 26, 2026.

- URL: https://blog.laozhang.ai/en/posts/free-ai-api-tiers-compared
- Published: 2026-07-02
- Updated: 2026-09-26
- Author: LaoZhang AI Team (https://blog.laozhang.ai/en/about)
- Category: API Guides
- Tags: Free AI API, Gemini API, Groq API, OpenRouter, Cloudflare Workers AI, LLM API, Rate Limits

---
For most developers, the best free AI API provider to start with is Google's Gemini API. A key from Google AI Studio gives you the Flash and Flash-Lite models, plus Gemini 2.5 Pro, at no token cost, without linking a billing account, and it works through an OpenAI-compatible endpoint. Two conditions change that answer. On the free tier, Google may use your prompts and outputs to improve its products, and human reviewers may read them, so it is the wrong place for customer data or confidential code. And the Gemini API is not offered in Russia, mainland China, or Hong Kong.

For a second key, Groq's Free Plan publishes exact limits (1,000 requests a day on `openai/gpt-oss-120b`), and OpenRouter's `:free` models put many models behind one key, capped at 50 requests a day until you have bought 10 credits. Three names you may expect are not free options anymore: Cerebras gives a $5 trial credit only after you add a payment method, GitHub Models was fully retired on July 30, 2026, and DeepSeek sells cheap tokens with no public free tier.

Pick by the constraint that matters most to you:

| If this matters most | Start with | Why | Watch for |
|---|---|---|---|
| Strongest free models, no card | Gemini API with an AI Studio key | Flash family and Gemini 2.5 Pro are free of charge on the free tier | Prompts may be used by Google; your limits are shown only in AI Studio |
| Prompts contain customer or confidential data | A paid tier, such as a Gemini project with billing linked | Paid Gemini content is not used to improve Google products | Keep free tiers for public or synthetic test data |
| Published, predictable limits | Groq Free Plan | 30 requests per minute and 1,000 per day on each listed chat model | Only 8,000 tokens per minute; short model list |
| Many models behind one key | OpenRouter `:free` models | One OpenAI-compatible endpoint for many model families | 50 requests per day unless you have bought 10 credits |
| Your app already runs on Cloudflare | Workers AI | 10,000 Neurons per day included, even on the Workers Free plan | Neurons are not tokens; several frontier models need paid billing |
| Your app serves users in the EEA, Switzerland, or the UK | A paid Gemini project or another paid API | Gemini's terms allow only Paid Services for apps offered there | You can test on free quota, but the launched app needs billing |
| You live in Russia, mainland China, or Hong Kong | Providers that list your country | Gemini, Claude, and OpenAI do not list these regions | Region rules are part of the terms, not a technical glitch |
| No request cap at all | An open-weight model running locally | No provider sets a daily limit | Your hardware and the model size it can run become the limit |

All limits here are as of September 26, 2026. Free tiers change often, and each provider's own console is the final word for your account.

## What "free" means depends on the provider

"Free AI API" covers four different arrangements, and the difference decides how long your project can lean on it.

A **recurring free tier** refills on a schedule. Gemini resets its daily request count at midnight Pacific time, Groq's Free Plan has daily request and token caps, and Cloudflare gives 10,000 Neurons every day. This is the only kind you can build a hobby project on for months.

A **one-time trial credit** is a small balance that runs out or expires. Cerebras gives $5 that expires after 30 days and only after you add a verified payment method. Anthropic gives new users a small amount of credit to test the Claude API, without a published figure. Hugging Face gives free accounts $0.10 a month for Inference Providers, which works as a sandbox rather than a workload.

**Free models on a router** are capped by the router, not by the model's maker. OpenRouter's `:free` variants follow OpenRouter's limits, and each free model is served by whichever upstream provider hosts it. A free Llama or Qwen model on OpenRouter says nothing about whether Meta or Alibaba has a free API.

**Local models** are the only setup with no provider-imposed cap. You download open weights and run them on your own machine, so the cost moves to hardware, electricity, and slower responses on smaller GPUs.

"Unlimited free AI API" offers from hosted services fall into none of these categories. When a hosted provider pays for GPUs, there is a cap somewhere, even if it is hidden behind "fair use" or a reseller's own quota.

![Four kinds of free AI API access compared: recurring free tiers, one-time trial credits, free models on a router, and local models, with how long each lasts](https://blog.laozhang.ai/posts/en/free-ai-api-tiers-compared/img/four-kinds-of-free.webp)

## Free AI API providers compared

As of September 26, 2026, from each provider's own pricing, rate-limit, or terms page:

| Provider | Kind of free | Card needed | Published free limit | Notes that change the choice |
|---|---|---|---|---|
| Google Gemini API | Recurring free tier | No, billing is only for paid tiers | Not published in the docs; per-model limits appear in AI Studio, per project | Free-tier content is used to improve Google products; not available in Russia, mainland China, or Hong Kong |
| Groq | Recurring Free Plan | Not listed as a requirement | `gpt-oss-120b`, `gpt-oss-20b`, `qwen3.8-27b`: 30 RPM, 1,000 RPD, 8,000 TPM, 200,000 tokens per day | Limits apply per organization, not per user |
| OpenRouter `:free` models | Router caps | No purchase needed for the base cap | 20 RPM and 50 requests per day; 1,000 per day once you have bought at least 10 credits | Extra accounts or keys do not add capacity; a negative balance can block free models |
| Cloudflare Workers AI | Recurring daily allocation | Not listed as a requirement | 10,000 Neurons per day, reset at 00:00 UTC | The listed Kimi, GLM-5.x, and DeepSeek V4 models need a paid billing method |
| Mistral | Free mode with included monthly usage | Check at signup | Amounts shown only on your account's Limits page | Pay-as-you-go extends the same account beyond the included usage |
| SambaNova Cloud | Free tier while no payment method is linked | No | DeepSeek-V3.1, Llama-3.3-70B, `gpt-oss-120b`: 20 RPM, 20 requests per day, 200,000 tokens per day | 20 requests a day is enough for testing only |
| NVIDIA API catalog | Free prototyping endpoints | NVIDIA Developer Program membership | No limits on the program page | Licensed for prototyping, research, and testing, not production |
| Hugging Face Inference Providers | Monthly credit | No | $0.10 of credit per month on a free account | Credit applies only when Hugging Face routes the request, not with your own provider key |
| Cohere | Trial key | Not listed as a requirement | 1,000 calls per month; chat at 20 requests per minute | Trial keys are for evaluation; production keys are paid |
| Cerebras | One-time trial credit | Yes, a verified payment method | $5 that expires after 30 days; 5 RPM, 1M tokens per day | No permanent free tier |
| GitHub Models | None | Not applicable | Retired July 30, 2026 | Playground, catalog, inference API, and bring-your-own-key access are gone |
| DeepSeek | None public | Paid top-up | Paid only: `deepseek-flash` from $0.15 per 1M input tokens off-peak | Cheap paid fallback rather than a free tier |
| Anthropic Claude API | Small starter credit | Paid after the credit | No published amount | No recurring free tier |
| OpenAI API | Complimentary tokens for eligible organizations that share traffic | Positive balance required | Depends on usage tier and model group | No universal signup credit |

RPM is requests per minute, RPD requests per day, and TPM tokens per minute. "Free of charge" on a pricing page means no token price, not unlimited use.

Image and video generation sit outside these free tiers. On the Gemini API, the Nano Banana image models and Veo show "Not available" for the free tier, and Batch and Flex processing are never free. For image APIs specifically, see [Free AI Image Generation API in 2026: What Is Actually Free?](https://blog.laozhang.ai/en/posts/free-ai-image-generation-api)

## Provider notes

### Gemini API: the strongest free models, with a data trade-off

Gemini's free tier is broad. As of September 26, 2026, the [pricing page](https://ai.google.dev/gemini-api/docs/pricing) lists input and output as "Free of charge" for Gemini 3.8 Flash (`gemini-3.8-flash`), 3.7 Flash, 3.6 Flash, 3.5 Flash and Flash-Lite, 3.1 Flash-Lite, 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite, Gemma 4, and Gemini Embedding 2. Gemini 3.1 Pro Preview is not free. Grounding with Google Search is free up to 500 requests a day, shared across 2.5 Flash and Flash-Lite, and not available on the free tier for Gemini 3.x models.

What Google does not publish anymore is a per-model table of free limits. The [rate-limit docs](https://ai.google.dev/gemini-api/docs/rate-limits) say limits are applied per project, not per API key, that the daily count resets at midnight Pacific time, and that the numbers for your project are shown in Google AI Studio. They also say capacity is not guaranteed. Any fixed "free Gemini RPM" figure you see elsewhere is either old or a guess. Creating more keys in the same project does not add quota, a point covered in detail in [Gemini API key free: where limits live and why more keys do not add quota](https://blog.laozhang.ai/en/posts/gemini-api-free-tier).

The trade-off is data. Under the [Gemini API Additional Terms](https://ai.google.dev/gemini-api/terms), content sent through unpaid quota is used to provide, improve, and develop Google products, human reviewers may read it, and Google asks you not to send sensitive, confidential, or personal information. The API becomes a Paid Service only when the Cloud project has an active billing account. Linking billing moves you to Tier 1 and changes the data terms, so it is the step to take before real user data goes through the API. Gemini 3.8 Flash costs $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026, rising to $1.50 and $7.50 from January 1, 2027.

Two region rules also apply. You can use the API only from a country on Google's available-regions list, which includes the United States, Japan, South Korea, Taiwan, and Spain but not Russia, mainland China, or Hong Kong. And if your app is offered to users in the EEA, Switzerland, or the UK, the terms allow only Paid Services for it.

### Groq: exact limits and fast responses on a short menu

Groq is the clearest free tier to plan against because the [rate-limit page](https://console.groq.com/docs/rate-limits) prints the numbers. On the Free Plan, `openai/gpt-oss-120b`, `openai/gpt-oss-20b`, and `qwen/qwen3.8-27b` each allow 30 requests per minute, 1,000 per day, 8,000 tokens per minute, and 200,000 tokens per day. Cached tokens do not count toward those limits. Limits belong to the organization, so inviting teammates does not multiply them.

The constraint you will hit first is usually 8,000 tokens per minute, not the request count. A single request with a 6,000-token prompt already uses most of that minute's budget. For chat-style prototypes with short prompts, the daily cap of 1,000 requests is the more realistic ceiling. The Free Plan chat list is short, with no Llama chat models in it as of September 26, 2026, so older tutorials that point to Llama IDs on Groq's free tier need updating.

### OpenRouter: one key, many free models, small daily cap

OpenRouter is useful when you do not know which model you want yet. Model IDs ending in `:free` cost nothing, and the [limits page](https://openrouter.ai/docs/api-reference/limits) sets the rules: 20 requests per minute, and 50 requests per day if you have bought fewer than 10 credits in total. Once your all-time purchases reach 10 credits, the daily cap rises to 1,000. That one-time purchase is a real cost.

Creating additional accounts or keys does not raise the cap, and a negative balance can cause errors even on free models. A 429 can come either from OpenRouter or from the upstream provider serving the free model, so a model that fails at busy hours may work later without any change on your side. The provider serving a free model also sets its own data policy, so read the model page before sending anything you would not publish. If your goal is a coding agent, [Claude Code with OpenRouter and DeepSeek](https://blog.laozhang.ai/en/posts/claude-code-openrouter-deepseek) walks through that setup.

### Cloudflare Workers AI: a daily allowance for Workers apps

[Workers AI](https://developers.cloudflare.com/workers-ai/platform/pricing/) is included in both the Free and Paid Workers plans, with 10,000 Neurons per day at no charge and a reset at 00:00 UTC. Neurons are Cloudflare's billing unit, and each model converts tokens to Neurons at its own rate, so check the model's row before estimating how many requests the allowance covers. Above the allowance you need Workers Paid, at $0.011 per 1,000 Neurons. Several of the strongest models, including `kimi-k2.6`, `glm-5.3`, and `deepseek-v4-pro-0813`, require a paid billing method even for small use. Workers AI fits best when your app already runs on Cloudflare and calls the model from a Worker.

### Mistral, SambaNova, NVIDIA, Hugging Face, and Cohere

These free tiers are real but narrower.

- **Mistral** Free mode lets you create API keys and use included monthly usage within the limits shown on your account's Limits page. Mistral does not publish those amounts, so open the Limits page after signup before you plan around it.
- **SambaNova Cloud** applies its Free tier whenever no payment method is linked: 20 requests per minute and 20 per day on DeepSeek-V3.1, Llama-3.3-70B, and `gpt-oss-120b`. That is enough to compare output quality, not to run an app.
- **NVIDIA** gives Developer Program members free access to NIM API endpoints on build.nvidia.com. NVIDIA describes that access as being for prototyping, research, development, and testing only, and the program page publishes no rate limits.
- **Hugging Face** gives free accounts $0.10 of Inference Providers credit per month. It is enough to confirm an integration works.
- **Cohere** trial keys allow 1,000 calls a month, with chat models at 20 requests per minute.

### No longer free, or never free: Cerebras, GitHub Models, DeepSeek, Claude, OpenAI

**Cerebras** answers "Is there a permanently free tier?" with "No" in its [rate-limit FAQ](https://inference-docs.cerebras.ai/support/rate-limits). New accounts get $5 of trial credit after adding a verified payment method, and the credit expires 30 days after it is granted. Without a payment method, the playground and API stay inactive.

**GitHub Models** was fully retired on July 30, 2026. GitHub's [documentation](https://docs.github.com/en/github-models/use-github-models/prototyping-with-ai-models) says the playground, model catalog, inference API, and bring-your-own-key access are no longer available to any customer, so tutorials built on it will fail.

**DeepSeek** has no public free tier. Its [pricing page](https://api-docs.deepseek.com/quick_start/pricing) lists only paid rates: `deepseek-flash` costs $0.15 per 1M input tokens (cache miss) and $0.60 per 1M output tokens off-peak, and twice that during peak hours (01:00–04:00 and 06:00–10:00 UTC on weekdays). That makes it a cheap first paid step rather than a free one.

**Anthropic** gives new users a small amount of credit to test the Claude API, with no recurring free tier after that. If you plan to keep using Claude, [How to Buy Claude API Access Safely](https://blog.laozhang.ai/en/posts/claude-api-key-free-tier) covers creating your own key and funding it. **OpenAI** has no universal signup credit; eligible organizations can receive daily complimentary tokens in exchange for sharing traffic, and the details are in [OpenAI API Free Tier: What's Actually Free and What Still Bills](https://blog.laozhang.ai/en/posts/openai-api-free-tier).

## Make the first request and handle a 429

Gemini, Groq, and OpenRouter all accept the OpenAI SDK with a different base URL, so one client can talk to all three. Keep each key in an environment variable and never paste keys into code or commit them.

A quick check from the terminal, using Groq:

```bash
curl https://api.groq.com/openai/v1/chat/completions \
  -H "Authorization: Bearer $GROQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "openai/gpt-oss-20b", "messages": [{"role": "user", "content": "Say hello in five words."}]}'
```

If you get JSON with a `choices` array, the key works. A 401 means the key is wrong or missing from the environment, and a 429 means you hit a rate limit.

### Where your live limits show up

| Provider | Where to see your current limit | When the daily count resets |
|---|---|---|
| Gemini API | Google AI Studio, per project | Midnight Pacific time |
| Groq | Account limits page, plus `x-ratelimit-remaining-requests` and `x-ratelimit-remaining-tokens` response headers; `retry-after` on a 429 | Shown in `x-ratelimit-reset-requests` |
| OpenRouter | `GET https://openrouter.ai/api/v1/key` returns `free_model_daily_requests` with used, limit, and remaining | UTC day |
| Cloudflare Workers AI | Workers AI dashboard, Neuron usage | 00:00 UTC |
| Mistral | Limits page in your account | Shown on that page |

For OpenRouter, the check is one line. The response reports the daily counter but not the per-minute limit.

```bash
curl https://openrouter.ai/api/v1/key -H "Authorization: Bearer $OPENROUTER_API_KEY"
```

A 429 on a per-minute limit clears within seconds. A 429 on a daily limit will not clear until the reset time above, so retrying in a loop only wastes time. That is where a second provider helps.

### A small fallback across free providers

This Python sketch tries Gemini first, then Groq, then an OpenRouter free model, and moves on only when a provider returns a 429. It skips any provider whose key is not set.

![Fallback flow that sends a prompt to Gemini, then Groq, then an OpenRouter free model on each 429, with per-minute and daily 429 handling](https://blog.laozhang.ai/posts/en/free-ai-api-tiers-compared/img/fallback-on-429.webp)

```python
import os
from openai import OpenAI, RateLimitError

PROVIDERS = [
    {
        "name": "gemini",
        "base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
        "key_env": "GEMINI_API_KEY",
        "model": "gemini-3.8-flash",
    },
    {
        "name": "groq",
        "base_url": "https://api.groq.com/openai/v1",
        "key_env": "GROQ_API_KEY",
        "model": "openai/gpt-oss-120b",
    },
    {
        "name": "openrouter",
        "base_url": "https://openrouter.ai/api/v1",
        "key_env": "OPENROUTER_API_KEY",
        # Pick any model ID ending in ":free" from openrouter.ai/models
        "model": os.environ.get("OPENROUTER_FREE_MODEL", ""),
    },
]


def ask(prompt: str) -> tuple[str, str]:
    last_error = None
    for p in PROVIDERS:
        key = os.environ.get(p["key_env"])
        if not key or not p["model"]:
            continue
        # One SDK retry covers short per-minute spikes; daily caps fall through.
        client = OpenAI(api_key=key, base_url=p["base_url"], max_retries=1)
        try:
            resp = client.chat.completions.create(
                model=p["model"],
                messages=[{"role": "user", "content": prompt}],
            )
            return p["name"], resp.choices[0].message.content
        except RateLimitError as err:
            last_error = err
    raise RuntimeError(f"Every configured provider is rate-limited: {last_error}")


if __name__ == "__main__":
    provider, text = ask("Explain what an API rate limit is in two sentences.")
    print(f"[{provider}] {text}")
```

Install the SDK with `pip install openai`, export `GEMINI_API_KEY`, `GROQ_API_KEY`, `OPENROUTER_API_KEY`, and `OPENROUTER_FREE_MODEL`, then run the file. Keep the order that matches your priorities. If your prompts should not go to Google's free tier, remove the Gemini entry rather than relying on it as a fallback.

Two cautions keep this pattern honest. The three models answer differently, so check that your prompts still work on each one before you depend on the fallback. And the fallback spreads your traffic across three sets of data terms, which is fine for test data and wrong for anything private.

## When free stops being the right plan

Free tiers are for learning, prototyping, and low-traffic personal tools. Move to a paid plan when any of these happens:

- **Real user data enters the prompts.** On Gemini, linking billing changes the data terms. On other providers, read the paid tier's data page before sending user content.
- **Your app will be offered to users in the EEA, Switzerland, or the UK** and uses Gemini. The terms require Paid Services there.
- **Daily caps shape the product.** If users see errors every evening because you ran out of requests, the free tier is costing you more than tokens would.
- **You need a model that is not free.** Gemini 3.1 Pro Preview, Gemini image generation, and several frontier models on Workers AI are paid-only.
- **You need stable capacity or support.** Gemini's docs say free capacity is not guaranteed, and NVIDIA's free endpoints are for prototyping only.

The next step depends on why you left. If you like the model you tested, add billing with the same provider: a Gemini project with billing, Groq's Developer Plan, or pay-as-you-go on Mistral keeps your code unchanged. If cost is the main worry, compare per-token prices first. DeepSeek's `deepseek-flash` at $0.15 per 1M input tokens off-peak is one low starting point, and [Cheapest LLM API Provider: Compare Price, Quality, Latency, and Gateway Risk](https://blog.laozhang.ai/en/posts/cheapest-llm-api-provider) lays out the wider comparison. If you would rather manage one OpenAI-compatible key and one prepaid balance across several model vendors, a paid gateway such as [laozhang.ai](https://laozhang.ai/) (base URL `https://api.laozhang.ai/v1`) is one option. It bills pay-as-you-go from a prepaid balance and does not pass on any provider's free quota.

## FAQ

### Is there an unlimited free AI API?

No hosted provider offers unlimited free use. Every free tier in the table above has a per-minute, daily, or monthly cap, and Gemini's docs state that even its free capacity is not guaranteed. The only setup without a provider cap is an open-weight model running on your own hardware.

### Which free AI APIs work without a credit card?

As of September 26, 2026, the Gemini API free tier works without a billing account, SambaNova's free tier applies precisely when no payment method is linked, and OpenRouter's `:free` models work at 50 requests a day without any purchase. Hugging Face's $0.10 monthly credit also comes with a free account. Cerebras is the notable exception: its trial credit requires a verified payment method.

### Do more API keys or accounts give me more free quota?

Not on the major providers. Gemini limits apply per project, Groq limits apply per organization, and OpenRouter states that additional accounts or keys do not change your limits.

### Is the DeepSeek API free?

No. DeepSeek publishes only paid rates, starting at $0.15 per 1M input tokens off-peak for `deepseek-flash`. It is inexpensive, which is why it appears in free API discussions, but you need a topped-up balance to call it.

### Are Cerebras and GitHub Models still free?

No. Cerebras has no permanent free tier, only a $5 trial credit that needs a payment method and expires after 30 days. GitHub Models was fully retired on July 30, 2026, and its inference API no longer accepts requests.

### Which free AI API is best for students?

For learning, start with a Gemini API key, which gives you capable models at no cost, and add Groq when you want published limits to design against. Use only practice data, since Gemini's free tier may use your prompts to improve Google products. If Gemini is not available in your country, check which providers list it before you sign up.
