← Back to articles
Unified API Calling

Cache Layer Design

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Background

Free quota is counted in RPM/RPD, but 30% of real traffic is "the same question asked by many users." For example, "Explain RAG in one sentence" might be asked a hundred times across users; hitting the free API a hundred times burns the daily RPD instantly. A cache layer intercepts this repetition, reserving free quota for genuinely new questions.

Two-Layer Cache Design

  • Exact cache (L1): hash(messages + model + temperature) → response, stored in Redis for 1 hour. Typical hit rate 20-40%, intercepting literal duplicates.
  • Semantic cache (L2): embed the prompt, store it in a vector store. New requests embed first, retrieve top-k similar prompts, and if cosine similarity > 0.92 and the cached response is still fresh, return it directly. This catches "same question, different wording."
  • Warm cache (L3, optional): pre-generate responses offline for high-frequency queries (product FAQ, doc summaries) and serve from a CDN — these never hit the gateway at all.

Hit Policy and Invalidation

  • Cacheability check: requests with temperature=0 and no tool_calls are cacheable; creative writing is not by default.
  • Tiered TTL: 24h for factual Q&A, 6h for code generation, 1h for time-sensitive content.
  • Active invalidation: when a model version upgrades, clear the corresponding prefix; users see x-cache: HIT/MISS in the response header.
  • Cache poisoning defense: prompts containing secrets (e.g. key=xxx) are never cached.

Code Example

import hashlib, redis
r = redis.Redis()
from openai import OpenAI
emb_client = OpenAI()

def cache_key(messages, model):
    return "v1:" + hashlib.sha256(
        (str(messages) + model).encode()).hexdigest()

def get_or_call(messages, model, call_fn):
    k = cache_key(messages, model)
    if cached := r.get(k):
        return cached
    # semantic check: embed and search neighbors
    emb = emb_client.embeddings.create(
        model="text-embedding-3-small",
        input=str(messages)).data[0].embedding
    if near := vector_db.search(emb, threshold=0.92):
        r.setex(k, 3600, near)
        return near
    resp = call_fn()
    r.setex(k, 3600, resp)
    vector_db.upsert(emb, resp)
    return resp

Embedding Model Choice

Semantic cache hit rate depends on the embedding model. OpenAI text-embedding-3-small (1536 dims) hits well but has cost; local bge-small-zh is free but slightly less accurate. Key trap: once you change the embedding model, the entire vector store becomes invalid and must be re-indexed. So the gateway must support "dual indexing" — the new model builds a new store while the old one keeps serving until its hit rate reaches zero, then is taken offline.

Best Practices

  • Measure hit rate: a rate below 10% usually means the cache key is too coarse or the TTL too short.
  • Bound cache size: LRU with a 1GB ceiling to avoid blowing up Redis.
  • Provide a bypass: for debugging, expose an x-bypass-cache: true header that skips the cache.
  • Cross-user isolation: never share cache entries across client_id boundaries to avoid sensitive data leaks.

Caching turns free quota into "unlimited" — as long as your question distribution has a long tail.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App — Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling — Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool — Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models — Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models — OpenAI Compatible + Smart Routing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.