Decision-focused list of free LLM APIs in 2026: which to pick, real limits, and how to stack free tiers so you never pay for inference.
2
5 commits
updated Jul 21, 2026
A curated, decision-focused guide to every LLM API you can use for $0 in 2026: which one to pick, what you can actually build, and how to stack them so you never pay for inference on a side project.
Most "free LLM API" lists are just a wall of provider names. This one answers the question you actually have: "I have a project and $0. Which do I use?"
⚠️ Free tiers change almost monthly. Every number below is a starting point, not a promise, so always confirm current limits on the provider's own docs before you build. Found something stale? Open a PR. That's what keeps this list worth starring.
| Your situation | Start here |
|---|---|
| General prototype, want a capable closed model | Google AI Studio (Gemini) |
| Need it fast (chat, voice, real-time) | Groq |
| Need high token volume / batch jobs | Cerebras or Mistral |
| Want lots of models through one key | OpenRouter |
| Already living inside GitHub | GitHub Models |
| Open-source maintainer | Claude for Open Source (see below) |
Golden rule: don't rely on one provider. Each has separate limits, so routing across a few multiplies your free capacity and gives you a fallback when one rate-limits you. See Stacking Strategy.
No expiry. Rate-limited, but genuinely usable for prototyping and low-volume apps.
The best all-round starting point. Gemini Flash on the free tier, no credit card, just a Google account. Roughly ~1,500 requests/day at a low per-minute cap, enough for most prototypes, and a genuinely capable model. Long context is a standout.
Runs open-weight models (Llama family, etc.) on custom LPU hardware, so it's very fast, often the lowest latency you'll find for free. OpenAI-compatible endpoints, so it drops into existing code. Limits vary by model (smaller models get much higher daily ceilings than the big ones).
Wafer-scale chips built for inference, with strong throughput and long-context handling. Generous daily token allowance on open-weight models. Great when you need volume without paying, especially batch work.
One of the most generous permanent quotas here (a large monthly token allowance), but the free "Experiment" tier requires opting into data training. Fine for hobby projects, think twice for anything sensitive.
A router that sits in front of 70+ providers. A subset of models is exposed as :free, and one key reaches many of the providers above. Free-model daily cap is modest and increases after a small one-time top-up. Unbeatable for trying many models quickly.
Join the free NVIDIA Developer Program and get starter inference credits with OpenAI-compatible endpoints. Hosts a big catalog of open models (DeepSeek, Llama, Qwen, GLM) and is often among the first to host new open-model drops.
Access a mix of models (including OpenAI and Llama) from inside GitHub. Experimentation-tier limits are modest. Best when your prompt-testing and comparison already happen where your code lives.
A free daily "neurons" allowance for inference at the edge. Ideal for small inference tasks in serverless apps already running on Cloudflare Workers.
A fixed amount of free usage. Great for evaluation, not for ongoing projects.
Launched in early 2026, Anthropic's Claude for Open Source program grants qualifying open-source maintainers several months of Claude Max at no cost, one of the largest free-access grants of the year, with limited spots. If you maintain an OSS project, it's worth checking eligibility.
The free tiers aren't competitors, they're lanes. A workload too chatty for one provider's per-minute limit fits comfortably in another's. A practical setup:
:free modelBecause each provider meters independently, routing across four can multiply your effective free capacity several times over, and it keeps your app alive if any one provider drops a model.
"Free" almost always has a cost that isn't dollars:
This list is only useful if it stays current, and it stays current because people like you fix it.
(Bonus: a merged PR here earns contributors GitHub's Pull Shark achievement.)
If this saved you time or money, star the repo. It helps other developers find it, and it's the only thanks an open list runs on.
MIT, free to use, share, and adapt.
Decision-focused list of free LLM APIs in 2026: which to pick, real limits, and how to stack free tiers so you never pay for inference.
2
5 commits
updated Jul 21, 2026
A curated, decision-focused guide to every LLM API you can use for $0 in 2026: which one to pick, what you can actually build, and how to stack them so you never pay for inference on a side project.
Most "free LLM API" lists are just a wall of provider names. This one answers the question you actually have: "I have a project and $0. Which do I use?"
⚠️ Free tiers change almost monthly. Every number below is a starting point, not a promise, so always confirm current limits on the provider's own docs before you build. Found something stale? Open a PR. That's what keeps this list worth starring.
| Your situation | Start here |
|---|---|
| General prototype, want a capable closed model | Google AI Studio (Gemini) |
| Need it fast (chat, voice, real-time) | Groq |
| Need high token volume / batch jobs | Cerebras or Mistral |
| Want lots of models through one key | OpenRouter |
| Already living inside GitHub | GitHub Models |
| Open-source maintainer | Claude for Open Source (see below) |
Golden rule: don't rely on one provider. Each has separate limits, so routing across a few multiplies your free capacity and gives you a fallback when one rate-limits you. See Stacking Strategy.
No expiry. Rate-limited, but genuinely usable for prototyping and low-volume apps.
The best all-round starting point. Gemini Flash on the free tier, no credit card, just a Google account. Roughly ~1,500 requests/day at a low per-minute cap, enough for most prototypes, and a genuinely capable model. Long context is a standout.
Runs open-weight models (Llama family, etc.) on custom LPU hardware, so it's very fast, often the lowest latency you'll find for free. OpenAI-compatible endpoints, so it drops into existing code. Limits vary by model (smaller models get much higher daily ceilings than the big ones).
Wafer-scale chips built for inference, with strong throughput and long-context handling. Generous daily token allowance on open-weight models. Great when you need volume without paying, especially batch work.
One of the most generous permanent quotas here (a large monthly token allowance), but the free "Experiment" tier requires opting into data training. Fine for hobby projects, think twice for anything sensitive.
A router that sits in front of 70+ providers. A subset of models is exposed as :free, and one key reaches many of the providers above. Free-model daily cap is modest and increases after a small one-time top-up. Unbeatable for trying many models quickly.
Join the free NVIDIA Developer Program and get starter inference credits with OpenAI-compatible endpoints. Hosts a big catalog of open models (DeepSeek, Llama, Qwen, GLM) and is often among the first to host new open-model drops.
Access a mix of models (including OpenAI and Llama) from inside GitHub. Experimentation-tier limits are modest. Best when your prompt-testing and comparison already happen where your code lives.
A free daily "neurons" allowance for inference at the edge. Ideal for small inference tasks in serverless apps already running on Cloudflare Workers.
A fixed amount of free usage. Great for evaluation, not for ongoing projects.
Launched in early 2026, Anthropic's Claude for Open Source program grants qualifying open-source maintainers several months of Claude Max at no cost, one of the largest free-access grants of the year, with limited spots. If you maintain an OSS project, it's worth checking eligibility.
The free tiers aren't competitors, they're lanes. A workload too chatty for one provider's per-minute limit fits comfortably in another's. A practical setup:
:free modelBecause each provider meters independently, routing across four can multiply your effective free capacity several times over, and it keeps your app alive if any one provider drops a model.
"Free" almost always has a cost that isn't dollars:
This list is only useful if it stays current, and it stays current because people like you fix it.
(Bonus: a merged PR here earns contributors GitHub's Pull Shark achievement.)
If this saved you time or money, star the repo. It helps other developers find it, and it's the only thanks an open list runs on.
MIT, free to use, share, and adapt.