⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Background
Model calls are inherently a black box: a single failure could be upstream 5xx, an expired key, an over-long prompt, a rate limit, or just DNS jitter. Without unified observability, ops relies on user complaints for diagnosis. Capturing logs, metrics, and traces at the gateway layer means every investigation starts from the same data set.
The Observability Triad
- Metrics: each request records
latency_p50/p95/p99,success_rate,tokens_in/out,cost_cents,fallback_count, sliced acrossclient_id×provider×model. - Logs: structured JSON, with fields for
request_id,client_id,provider,upstream_latency,error_class,fallback_chain. Retain 7 days of raw logs + 90 days of aggregates. - Traces: OpenTelemetry spans cover "client → gateway → upstream → parse → return," so a single failure pinpoints its exact segment in seconds.
Key Dashboard Panels
- Health quadrants: latency (upstream + gateway), success rate, cost (daily avg + monthly cumulative), remaining quota. Each has an SLO and alert threshold.
- Provider heatmap: success-rate distribution per provider across time slots, used to tune routing weights.
- Prompt length histogram: long prompts risk blowing the context window — catch "user input inflation" early.
Code Example
from opentelemetry import trace, metrics
tracer = trace.get_tracer("gateway")
meter = metrics.get_meter("gateway")
latency_hist = meter.create_histogram("gateway.latency", unit="ms")
token_counter = meter.create_counter("gateway.tokens")
async def call_upstream(req, provider):
with tracer.start_as_current_span(f"upstream.{provider}") as span:
span.set_attribute("model", req["model"])
t0 = time.time()
try:
resp = await http.post(...)
latency_hist.record((time.time()-t0)*1000, {
"provider": provider, "model": req["model"]})
token_counter.add(resp["usage"]["total_tokens"], {
"provider": provider, "dir": "total"})
return resp
except Exception as e:
span.record_exception(e)
span.set_status(trace.Status(trace.StatusCode.ERROR))
raise
Cost vs Observability Balance
Capturing every trace blows up storage; skipping entirely loses debuggability. The balance is tiered sampling: 1% of normal requests, 100% of error requests, 100% of slow requests (above P99 threshold). 99% of traffic never hits disk, but every "worth-seeing" request does. Additionally, tier retention: 7 days raw logs, 90 days aggregates, 1 year cold archive — query performance and storage cost both optimized.
Best Practices
- Sampling: full trace capture explodes storage; sample 1% randomly plus 100% of error requests.
- PII redaction: prompts often contain user PII; scrub email/phone from logs before they hit disk.
- SLO burn rate: use multi-window multi-burn-rate alerting (5m × 1h + 1h × 6h) to catch anomalies earlier than static thresholds.
- Unified ID threading: every request carries matching trace_id / span_id / log_id for cross-table queries.
Observability is not "investigate after the incident" — it is "see the system even when nothing is wrong."
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key