โ† Back to articles
Unified API Calling

Retry and Fallback Strategy

โš ๏ธ Pending Update ยท 2026-08-29 Verification ยท Content may be outdated, please refer to official docs Updated: 2026-08-29 ยท Status: Pending Verification

Background

Free models are inherently less stable than paid: OpenRouter free models frequently 429, Groq returns 5xx during peak, Hugging Face cold starts can take 30s. If clients call them directly, every call site ends up implementing its own retry and fallback, which is both repetitive and error-prone. Concentrating these strategies in the gateway lets clients care only about the business result.

Core Strategies

  • Error taxonomy: the gateway sorts upstream errors into four buckets โ€” retryable (429, 503, timeout), non-retryable (400, 401), degradable (model lacks tool support), fatal (sustained 5xx). Each bucket follows a different path.
  • Exponential backoff with jitter: retry delay = base * 2^attempt + random(0,1) to avoid synchronized retries hammering the upstream. Base 200ms, max 4 attempts, total under 8s.
  • Cross-provider fallback: the same model often has replicas across providers (e.g. llama-3.3-70b on both Groq and OpenRouter). On first failure, switch providers instead of beating the same one.
  • Circuit breaker: if a provider's failure rate crosses 50% in a 5-minute window, break for 60s and route around it.
  • Degradation chain: user requests gpt-4o โ†’ free equivalent deepseek-chat:free โ†’ cheaper llama-3.1-8b:free โ†’ reject. Each hop is recorded in response headers so clients can observe it.

Code Example

import asyncio, random

async def call_with_fallback(call_fn, providers, max_attempts=4):
    for attempt in range(max_attempts):
        provider = providers[attempt % len(providers)]
        try:
            return await call_fn(provider)
        except (TimeoutError, RateLimitError):
            # exponential backoff + jitter
            delay = 0.2 * (2 ** attempt) + random.random()
            await asyncio.sleep(delay)
        except (InternalError, ServiceUnavailable):
            # immediately switch provider, no retry on same one
            continue
    # all attempts failed: degrade to cheaper model or raise
    raise FallbackExhausted("tried: " + ",".join(providers))

Partial Retry Semantics

For non-streaming multi-turn tool-call scenarios, retry semantics get tricky: round one succeeded and called a tool, round two failed. If you retry the whole thing, the tool's side effect fires twice. The right approach is an idempotency token: the client generates a unique ID for the whole session, and the gateway passes "already-succeeded tool_call_ids" as a passthrough header so the upstream skips executed parts. Tool calls without an idempotency token allow retries only on the first round; later failures must fail the whole chain.

Best Practices

  • Idempotency required: retries are safe only for idempotent requests (pure generation); retries on tool_calls with side effects must be more conservative.
  • Never retry streaming: once a stream has emitted tokens it cannot be retried โ€” only pre-first-chunk failures may retry.
  • Retry budget: cap maximum tokens spent on retries, so a single failure doesn't trigger four retries that double your burn.
  • Observability: tag every fallback with a fallback_chain metric label so ops can spot the most-traveled paths.
  • Circuit-breaker isolation: each provider gets its own breaker so one's outage never takes down another.

The question with free models is never "can it be used" but "who catches it when it breaks," and the answer is always the gateway.

๐Ÿš€ Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key โ€” one key, 100+ models, free models at zero cost.

๐Ÿ‘‰ Register on Apishare.cc โ†’ Get your unified API Key

๐Ÿ“Š Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings โ†’


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App โ€” Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling โ€” Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool โ€” Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models โ€” Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models โ€” OpenAI Compatible + Smart Routing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide โ€” sign up and get bonus credits.