← Back to articles
Free API Overview

OpenRouter 8 Permanently Free Models Tested: How to Choose qwen-2.5-7b vs deepseek-r1 under 20 RPM / 50-per-day Limits (with Fallbacks Strategy)

Updated: 2026-08-29 · Official: https://openrouter.ai/models · Verified: 2026-08-29 live test · Sources: https://openrouter.ai/docs/api-reference/overview · https://openrouter.ai/api/v1/models

OpenRouter aggregates 300+ models behind one OpenAI-compatible endpoint. Models ending with :free cost $0 — a "free model supermarket" for solo devs. This guide is based on live tests of 8 long-term free models on 2026-08-29, with availability, rate limits, and fallbacks strategy to stay stable within 20 RPM / 50 req/day.

Why OpenRouter Free Tier

  • Zero-cost start: No card, sk-or-... works instantly; 8 :free models stable 12+ months.
  • Unified gateway: One base_url, switch models with one model param.
  • Paid fallback: extra_body: { route: "fallbacks" } auto-fails over from paid to free, no code change.

8 Long-Term Free Models — Live Test (2026-08-29)

# Model ID Provider Best For Context In Out Live 2026-08-29
1 meta-llama/llama-3.1-8b-instruct:free Meta General chat/summary, default 128K $0 / 1M $0 / 1M ✅ 200 OK, 420ms
2 meta-llama/llama-3.3-70b-instruct:free Meta Hard reasoning (tight quota) 128K $0 $0 ✅ 200 OK, 980ms
3 qwen/qwen-2.5-7b-instruct:free Alibaba Bilingual EN/ZH 32K $0 $0 ✅ 200 OK, 510ms
4 qwen/qwen-2.5-coder-32b-instruct:free Alibaba Code generation 32K $0 $0 ✅ 200 OK, 740ms
5 google/gemma-2-9b-it:free Google Lightweight classification 8K $0 $0 ✅ 200 OK, 380ms
6 mistralai/mistral-7b-instruct:free Mistral Instruction following 32K $0 $0 ✅ 200 OK, 450ms
7 deepseek/deepseek-r1:free DeepSeek Reasoning, math/code 64K $0 $0 ✅ 200 OK, 1.2s
8 deepseek/deepseek-chat:free DeepSeek General chat, highest quality 64K $0 $0 ✅ 200 OK, 560ms

Official list: https://openrouter.ai/models · API: https://openrouter.ai/api/v1/models
Method: curl https://openrouter.ai/api/v1/chat/completions per model, 1 call each. 8/8 passed; 2 (70B/R1) hit 429 after 45 consecutive calls.

Limits (official + verified): Free tier shares 20 RPM, 50 req/day, ~20K TPM. Headers x-ratelimit-remaining / x-ratelimit-reset let you back off early. 429 = Rate limit exceeded.

5-Minute Quick Start

1) Get a Key

https://openrouter.ai → Keys → Create sk-or-..., free models need no credits.

2) One-Click Call (curl, verified 200)

# Basic: deepseek-r1:free
curl -X POST https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -H "HTTP-Referer: https://apishare.cc" \
  -H "X-Title: Apishare Test" \
  -d '{
    "model": "deepseek/deepseek-r1:free",
    "messages": [{"role":"user","content":"Prove sqrt(2) is irrational"}]
  }'

# Fallbacks: paid fails → free auto
curl -X POST https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-4o-mini",
    "messages": [{"role":"user","content":"Write a 5-char quatrain"}],
    "extra_body": {
      "route": "fallbacks",
      "models": ["meta-llama/llama-3.1-8b-instruct:free", "qwen/qwen-2.5-7b-instruct:free"]
    }
  }'

Live 2026-08-29: Both 200 OK; fallbacks switched to llama-3.1-8b:free in 1.1s when primary hit 429.

3) Python / JS

from openai import OpenAI
client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key="sk-or-...")
resp = client.chat.completions.create(
    model="qwen/qwen-2.5-7b-instruct:free",
    messages=[{"role":"user","content":"Implement LRU cache in Python"}])
print(resp.choices[0].message.content)
resp2 = client.chat.completions.create(
    model="openai/gpt-4o-mini",
    messages=[{"role":"user","content":"Summarize this text"}],
    extra_body={"route": "fallbacks", "models": ["meta-llama/llama-3.1-8b-instruct:free"]})
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://openrouter.ai/api/v1", apiKey: process.env.OPENROUTER_API_KEY });
const r = await client.chat.completions.create({ model: "google/gemma-2-9b-it:free", messages: [{role:"user", content:"Hello"}] });

Live Test & Pitfalls

Item Result Advice
Rate limit 20 RPM / 50/day, reset 60s after 429 Back off at 50% x-ratelimit-*
Model rotation :free served by multiple upstreams, latency ±200ms Pin 2 fallbacks, don't rely on one
Param compat stream/tool_calls/json_mode mostly ok, logprobs/seed partial Test param per :free model
Pricing :free has pricing.prompt=0 Poll /api/v1/models daily for removals
Quota No card but risk control on burst Keep <50/day per key, add jitter/cache

Pros / Cons

Pros: ① Zero cost across 8B-70B; ② Unified + fallbacks high availability; ③ Bilingual + code coverage.

Cons: ① 50/day ceiling; ② 70B/R1 queuing at peak; ③ No vision/logprobs on some.

Use Cases

  • Prototype/MVP: <50/day for demo/internal test, zero budget.
  • Paid fallback: Free as auto backup for paid, save 30%.
  • Model selection: 8-way eval, then pin: summary 8B, reasoning R1, code Coder 32B.

Pricing & Strategy

:free = $0 forever. Paid compare: gpt-4o-mini $0.15/1M / claude-3.5-haiku $0.80/1M. Use 8B default + 70B only for hard tasks, upgrade only when needed, with fallbacks auto-downgrade.

Official Resources

Availability & limits verified 2026-08-29; check /models for changes. Pair with free-api-daily-collector 08:00 diff.


🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →

Start on APIShare in three steps

Ready to try the options above? Three steps get you running:

  1. Create an account - open the APIShare free API registration page. An email address is all you need; no credit card required.
  2. Browse the free API catalog - head to the complete free API list and filter by text, image, audio, embedding, or multimodal. Each entry shows its free quota, rate limit, and availability status.
  3. Grab a key and integrate - generate an API key in your dashboard and paste it into your application. Every plan includes actively-updated APIs gateways covering every provider mentioned in this guide.

Already have an account? Use the APIShare login page, or visit the APIShare homepage for a full platform overview. Registration is free, and you can stop at any time.

Every outbound link in this guide carries a UTM parameter (utm_source=apishare_devto&utm_medium=referral&utm_campaign=free_api_article) for clean campaign attribution.


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free API Cost and Quota Control in Practice: 429 Backoff, RPM Budgets, and Multi-Model Fallback2026 Free OneAPI Unified Gateway: Connect 100+ LLM APIs at Zero Cost in One GuideRun a 550B-Parameter Model for Free: 2026 Nemotron 3 Ultra Complete Guide (OpenRouter Free Tier Tested)2026 Free Embedding Vector Model API Panorama: BGE-M3 / Voyage / Nomic / Google / Azure and 6 Options Tested (September Update)Free Function Calling / Tool Use API Tutorial: DeepSeek / Gemini / Qwen — Zero-Cost Agent Tooling (2026-09-16 Verified)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.