โ† Back to articles
Unified API Calling

Build Your Own Unified API Proxy

โš ๏ธ Pending Update ยท 2026-08-29 Verification ยท Content may be outdated, please refer to official docs Updated: 2026-08-29 ยท Status: Pending Verification

Background

After reading the previous 14 articles, you probably want to try building one yourself. This article gives a minimal-viable gateway in under 100 lines: it accepts OpenAI-shape requests, routes by the model field to any OpenAI-compatible upstream, supports streaming passthrough, and does simple fallback. Before production use, add auth, rate limiting, and metrics โ€” but as a prototype it is clear enough.

Minimal Architecture

  • Ingress: POST /v1/chat/completions, accepts the standard OpenAI request body.
  • Routing table: a YAML config file mapping model -> (base_url, api_key, fallback).
  • Client: a pooled httpx.AsyncClient.
  • Streaming: pass through text/event-stream chunk by chunk.
  • Fallback: on primary 5xx, switch to the fallback provider.

Code Example

# gateway.py - minimal viable unified API proxy
from fastapi import FastAPI, Request, HTTPException
from fastapi.responses import StreamingResponse
import httpx, yaml

app = FastAPI()
cfg = yaml.safe_load(open("providers.yaml"))
CLIENT = httpx.AsyncClient(timeout=120)

@app.post("/v1/chat/completions")
async def chat(req: Request):
    body = await req.json()
    model = body.get("model", "")
    route = next((r for r in cfg["routes"] if r["model"] == model), None)
    if not route:
        raise HTTPException(404, f"model {model} not configured")
    headers = {"Authorization": f"Bearer {route['key']}",
               "Content-Type": "application/json"}
    # streaming passthrough
    if body.get("stream"):
        async def gen():
            async with CLIENT.stream("POST",
                    f"{route['base']}/chat/completions",
                    json=body, headers=headers) as r:
                async for line in r.aiter_lines():
                    yield line + "\n"
        return StreamingResponse(gen(), media_type="text/event-stream")
    # non-streaming with fallback
    try:
        r = await CLIENT.post(f"{route['base']}/chat/completions",
                              json=body, headers=headers)
        return r.json()
    except (httpx.HTTPError, httpx.TimeoutException):
        if route.get("fallback"):
            body["model"] = route["fallback"]
            return await chat(req)
        raise HTTPException(502, "upstream failed")

Companion providers.yaml:

routes:
  - model: groq/llama-3.3-70b
    base: https://api.groq.com/openai/v1
    key: gsk_xxx
    fallback: openrouter/llama-3.3-70b:free
  - model: openrouter/llama-3.3-70b:free
    base: https://openrouter.ai/api/v1
    key: sk-or-xxx

Launch with uvicorn gateway:app --port 8000, point the client's base_url to http://localhost:8000/v1, and you are ready.

Production Hardening Checklist

After the minimal gateway ships, pushing it to "production-ready" requires seven additions: (1) auth โ€” JWT or API key validation; (2) rate limiting โ€” token bucket per client_id; (3) metrics โ€” Prometheus /metrics endpoint; (4) logs โ€” structured JSON with PII scrubbing; (5) tracing โ€” OpenTelemetry trace; (6) cache โ€” Redis exact layer; (7) health checks โ€” /healthz and /readyz separated. Complete the first six and the gateway handles around 1000 QPS; the seventh is mandatory for Kubernetes deployment.

Best Practices and Next Steps

  • Add auth: use FastAPI Depends to validate the client token, otherwise your gateway will be free-ridden.
  • Add rate limiting: slowapi or a self-implemented token bucket, 30 calls/min per client_id.
  • Add metrics: a Prometheus /metrics endpoint for latency and success rate.
  • Add a cache: a Redis exact-cache layer doubles hit rate immediately.
  • Add a dashboard: a Grafana template gets ops experience close to a commercial gateway.
  • Load test: before going live, run wrk or vegeta to learn your peak QPS and P99 latency.
  • Regression tests: after every change, run integration tests covering "call โ†’ stream โ†’ fallback" so behavior never regresses silently.

Minimal viable gateway โ†’ add auth โ†’ add rate limiting โ†’ add metrics โ†’ add cache โ†’ add degradation chain. Each step lands one of the concepts from articles 1-14 in code.

๐Ÿš€ Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key โ€” one key, 100+ models, free models at zero cost.

๐Ÿ‘‰ Register on Apishare.cc โ†’ Get your unified API Key

๐Ÿ“Š Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings โ†’


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App โ€” Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling โ€” Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool โ€” Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models โ€” Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models โ€” OpenAI Compatible + Smart Routing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide โ€” sign up and get bonus credits.