⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Introduction
Almost every free API has rate limits, in three common forms: RPM (requests per minute), TPM (tokens per minute), and daily quota (requests per day). This article explains how to identify and handle them so your scripts do not get banned, with full code for backoff retry and multi-provider failover.
架构图
Identify Rate-limit Responses
Each platform returns 429 differently, but key info is always present:
- OpenAI / OpenRouter:
headers["x-ratelimit-remaining-requests"] - Groq:
error.code == "rate_limit_exceeded" - Gemini:
error.status == 429withRESOURCE_EXHAUSTED
The correct approach is to catch RateLimitError rather than only checking HTTP status. Also log x-ratelimit-* headers for early warning.
Exponential Backoff Retry
import os, time, random
from openai import OpenAI, RateLimitError
client = OpenAI(
api_key=os.environ["GROQ_API_KEY"],
base_url="https://api.groq.com/openai/v1",
)
def chat_with_retry(messages, max_retries=5):
for attempt in range(max_retries):
try:
resp = client.chat.completions.create(
model="llama-3.3-70b-versatile",
messages=messages,
max_tokens=256,
)
return resp.choices[0].message.content
except RateLimitError as e:
if attempt == max_retries - 1:
raise
# Exponential backoff + jitter to avoid thundering herd
wait = (2 ** attempt) + random.uniform(0, 1)
print(f"rate limited, retry in {wait:.1f}s")
time.sleep(wait)
print(chat_with_retry([{"role":"user","content":"hi"}]))
Multi-provider Failover
When one provider hits the limit, switch to another:
PROVIDERS = [
("groq", OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1"), "llama-3.3-70b-versatile"),
("together", OpenAI(api_key=os.environ["TOGETHER_API_KEY"], base_url="https://api.together.xyz/v1"), "meta-llama/Llama-3.3-70B-Instruct-Turbo"),
("openrouter",OpenAI(api_key=os.environ["OPENROUTER_API_KEY"],base_url="https://openrouter.ai/api/v1"), "meta-llama/llama-3.3-70b-instruct:free"),
]
def chat_failover(prompt):
for name, client, model in PROVIDERS:
try:
r = client.chat.completions.create(
model=model,
messages=[{"role":"user","content":prompt}],
max_tokens=128,
)
return name, r.choices[0].message.content
except Exception as e:
print(f"[{name}] failed: {e}, trying next")
raise RuntimeError("all providers failed")
who, ans = chat_failover("hi")
print(f"answered by {who}: {ans}")
Local Token Bucket
Throttle outbound requests proactively to avoid server-side rejection:
import time, threading
class TokenBucket:
def __init__(self, rate, capacity):
self.rate = rate
self.capacity = capacity
self.tokens = capacity
self.last = time.monotonic()
self.lock = threading.Lock()
def acquire(self):
with self.lock:
now = time.monotonic()
self.tokens = min(self.capacity, self.tokens + (now - self.last) * self.rate)
self.last = now
if self.tokens < 1:
wait = (1 - self.tokens) / self.rate
time.sleep(wait)
self.tokens = 0
else:
self.tokens -= 1
bucket = TokenBucket(rate=15/60, capacity=5) # 15 RPM, burst 5
def safe_chat(prompt):
bucket.acquire()
return chat_with_retry([{"role":"user","content":prompt}])
Practical Tips
- Local cache: Cache results keyed by prompt hash for direct hits.
- Batching: Merge independent requests into a single batch (supported by OpenRouter).
- Off-peak: Free tiers are busiest at peak hours; schedule heavy jobs at midnight.
- Monitor quota: Log
x-ratelimit-remainingfor early warning.
Monitoring and Alerting
In production, log every request's x-ratelimit-remaining-* header to a time-series DB (e.g. Prometheus + Grafana) and alert when remaining quota drops below 20%. Also record each 429 with the provider and timestamp to pinpoint which one hit the cap. A simple approach is to add a decorator around chat_with_retry that ships exceptions and retry counts to your log service. TokenBucket is already locked, so multiple threads can share one instance; cache stampedes are solved by request coalescing (one in-flight call per prompt within 100ms).
Troubleshooting
- Still 429 after backoff: Likely daily quota exhausted. Wait until tomorrow or switch providers.
- TPM exceeded: Reduce
max_tokensor trim the message history. - IP banned: Providers share anti-abuse signals. Always include backoff.
Good rate-limit handling multiplies the value of your free credits tenfold.
Backoff Algorithm Example
import time, random
def retry_with_backoff(call, max_retries=5, base_delay=1.0):
for i in range(max_retries):
try:
return call()
except (RateLimitError, ServerError) as e:
if i == max_retries - 1: raise
delay = base_delay * (2 ** i) + random.uniform(0, 1)
time.sleep(delay)
Best Practices
- Exponential backoff + jitter: pure exponential makes multiple clients retry in sync; add jitter to avoid the thundering herd.
- Cap backoff at 60s: a single backoff over 60s means the user has already left; further backoff is pointless.
- Switch providers after 2 consecutive 429s: do not stick with one provider once it starts throttling.
- Identify limit type from headers: the
X-RateLimit-Remainingresponse header tells you how much is left, distinguishing RPM vs daily quota.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key