⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Background
A single provider cannot sustain high concurrency: Groq free tier RPM=30, OpenRouter free 20 RPM. But if you register five providers simultaneously, the theoretical total of 100 RPM is plenty for a small product. The load balancer exposes these five as one "logical provider," distributing internal traffic by policy.
Load-Balancing Strategies
- Weighted Round-Robin: each provider's weight = remaining free quota / average latency; the more generous and faster, the higher the weight. Simple and stable — the default.
- Least Connections: tracks each provider's in-flight requests and prefers the idlest. Suits long-lived / streaming scenarios.
- Latency-aware (P2C + EWMA): Pick-of-2 candidates, choose the one with lower EWMA latency. Originated in Nginx, sensitive to latency jitter.
- Quota-aware: skip providers whose daily quota is exhausted, so retries are not wasted.
- Sticky: route the same session to the same provider to avoid context-window drift caused by differing token counters.
Code Example
import time, random
from collections import defaultdict
class LoadBalancer:
def __init__(self, providers):
# provider -> {"weight": int, "ewma_latency": float, "inflight": int}
self.p = {k: {"weight": v, "ewma_latency": 100, "inflight": 0}
for k, v in providers.items()}
def pick(self) -> str:
# Pick-of-2 by latency, weighted fallback
candidates = random.sample(list(self.p.keys()), 2)
candidates = [c for c in candidates if self.p[c]["inflight"] < 5]
if not candidates:
return min(self.p, key=lambda k: self.p[k]["inflight"])
return min(candidates, key=lambda k: self.p[k]["ewma_latency"])
def observe(self, provider, latency, ok):
a = 0.3 # EWMA smoothing
self.p[provider]["ewma_latency"] = (
(1-a) * self.p[provider]["ewma_latency"] + a * latency)
if not ok:
self.p[provider]["weight"] = max(1, self.p[provider]["weight"] // 2)
Health Check Cadence
Background health check cadence is critical: too frequent (per second) burns provider quota, too sparse (per minute) lags detection. Recommended: 5-second light probes (hit /models), 15-second heavy probes (send one cheapest real prompt). Light probes test connectivity, heavy probes test inference availability — together they distinguish "network up but service down" from "fully offline." Heavy probes cost provider quota, so run them only on idle providers with no traffic.
Trade-offs and Best Practices
- Avoid pure random: random distribution causes occasional single-provider spikes; weighted + P2C is more stable.
- Async health checks: ping
/modelsevery 5s in the background — never make real requests bear the probing cost. - Warm up new providers: a newly added provider starts at 5% weight, not 100%, ramping up to avoid cold-start stampedes.
- Sticky sessions: route the same session to the same provider to avoid context-window drift.
- Pool cap protection: set a max-concurrency per provider so one cannot be over-stamped.
Load balancing is not "fairness" — it is "keeping each provider inside its comfort zone."
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key