← Back to articles
Detailed Usage

Rate Limits of Free APIs and How to Handle Them

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Introduction

Almost every free API has rate limits, in three common forms: RPM (requests per minute), TPM (tokens per minute), and daily quota (requests per day). This article explains how to identify and handle them so your scripts do not get banned, with full code for backoff retry and multi-provider failover.

架构图

flowchart TD A[429 Too Many Requests] --> B[Exponential backoff] B --> C[Retry after delay] C -->|Still 429| D[Switch to next provider] D --> A C -->|Success| E[Response returned] A --> F[Identify limit type] F -->|RPM| G[Spread calls in time] F -->|TPM| H[Shorten prompts] F -->|Daily quota| I[Switch provider]

Identify Rate-limit Responses

Each platform returns 429 differently, but key info is always present:

  • OpenAI / OpenRouter: headers["x-ratelimit-remaining-requests"]
  • Groq: error.code == "rate_limit_exceeded"
  • Gemini: error.status == 429 with RESOURCE_EXHAUSTED

The correct approach is to catch RateLimitError rather than only checking HTTP status. Also log x-ratelimit-* headers for early warning.

Exponential Backoff Retry

import os, time, random
from openai import OpenAI, RateLimitError

client = OpenAI(
    api_key=os.environ["GROQ_API_KEY"],
    base_url="https://api.groq.com/openai/v1",
)

def chat_with_retry(messages, max_retries=5):
    for attempt in range(max_retries):
        try:
            resp = client.chat.completions.create(
                model="llama-3.3-70b-versatile",
                messages=messages,
                max_tokens=256,
            )
            return resp.choices[0].message.content
        except RateLimitError as e:
            if attempt == max_retries - 1:
                raise
            # Exponential backoff + jitter to avoid thundering herd
            wait = (2 ** attempt) + random.uniform(0, 1)
            print(f"rate limited, retry in {wait:.1f}s")
            time.sleep(wait)

print(chat_with_retry([{"role":"user","content":"hi"}]))

Multi-provider Failover

When one provider hits the limit, switch to another:

PROVIDERS = [
    ("groq",      OpenAI(api_key=os.environ["GROQ_API_KEY"],      base_url="https://api.groq.com/openai/v1"),       "llama-3.3-70b-versatile"),
    ("together",  OpenAI(api_key=os.environ["TOGETHER_API_KEY"],  base_url="https://api.together.xyz/v1"),          "meta-llama/Llama-3.3-70B-Instruct-Turbo"),
    ("openrouter",OpenAI(api_key=os.environ["OPENROUTER_API_KEY"],base_url="https://openrouter.ai/api/v1"),         "meta-llama/llama-3.3-70b-instruct:free"),
]

def chat_failover(prompt):
    for name, client, model in PROVIDERS:
        try:
            r = client.chat.completions.create(
                model=model,
                messages=[{"role":"user","content":prompt}],
                max_tokens=128,
            )
            return name, r.choices[0].message.content
        except Exception as e:
            print(f"[{name}] failed: {e}, trying next")
    raise RuntimeError("all providers failed")

who, ans = chat_failover("hi")
print(f"answered by {who}: {ans}")

Local Token Bucket

Throttle outbound requests proactively to avoid server-side rejection:

import time, threading

class TokenBucket:
    def __init__(self, rate, capacity):
        self.rate = rate
        self.capacity = capacity
        self.tokens = capacity
        self.last = time.monotonic()
        self.lock = threading.Lock()
    def acquire(self):
        with self.lock:
            now = time.monotonic()
            self.tokens = min(self.capacity, self.tokens + (now - self.last) * self.rate)
            self.last = now
            if self.tokens < 1:
                wait = (1 - self.tokens) / self.rate
                time.sleep(wait)
                self.tokens = 0
            else:
                self.tokens -= 1

bucket = TokenBucket(rate=15/60, capacity=5)  # 15 RPM, burst 5
def safe_chat(prompt):
    bucket.acquire()
    return chat_with_retry([{"role":"user","content":prompt}])

Practical Tips

  • Local cache: Cache results keyed by prompt hash for direct hits.
  • Batching: Merge independent requests into a single batch (supported by OpenRouter).
  • Off-peak: Free tiers are busiest at peak hours; schedule heavy jobs at midnight.
  • Monitor quota: Log x-ratelimit-remaining for early warning.

Monitoring and Alerting

In production, log every request's x-ratelimit-remaining-* header to a time-series DB (e.g. Prometheus + Grafana) and alert when remaining quota drops below 20%. Also record each 429 with the provider and timestamp to pinpoint which one hit the cap. A simple approach is to add a decorator around chat_with_retry that ships exceptions and retry counts to your log service. TokenBucket is already locked, so multiple threads can share one instance; cache stampedes are solved by request coalescing (one in-flight call per prompt within 100ms).

Troubleshooting

  • Still 429 after backoff: Likely daily quota exhausted. Wait until tomorrow or switch providers.
  • TPM exceeded: Reduce max_tokens or trim the message history.
  • IP banned: Providers share anti-abuse signals. Always include backoff.

Good rate-limit handling multiplies the value of your free credits tenfold.

Backoff Algorithm Example

import time, random

def retry_with_backoff(call, max_retries=5, base_delay=1.0):
    for i in range(max_retries):
        try:
            return call()
        except (RateLimitError, ServerError) as e:
            if i == max_retries - 1: raise
            delay = base_delay * (2 ** i) + random.uniform(0, 1)
            time.sleep(delay)

Best Practices

  • Exponential backoff + jitter: pure exponential makes multiple clients retry in sync; add jitter to avoid the thundering herd.
  • Cap backoff at 60s: a single backoff over 60s means the user has already left; further backoff is pointless.
  • Switch providers after 2 consecutive 429s: do not stick with one provider once it starts throttling.
  • Identify limit type from headers: the X-RateLimit-Remaining response header tells you how much is left, distinguishing RPM vs daily quota.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free AI Content Moderation API Guide 2026: Llama Guard 3 vs Perspective vs OpenAIFree OCR and Document Parsing API in PracticeIntegrating Free APIs into Your Local IDEConnecting Free Models to OpenCode in PracticeApplying for an OpenRouter API Key and Understanding Pricing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.