⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Background
Free quota is counted in RPM/RPD, but 30% of real traffic is "the same question asked by many users." For example, "Explain RAG in one sentence" might be asked a hundred times across users; hitting the free API a hundred times burns the daily RPD instantly. A cache layer intercepts this repetition, reserving free quota for genuinely new questions.
Two-Layer Cache Design
- Exact cache (L1):
hash(messages + model + temperature) → response, stored in Redis for 1 hour. Typical hit rate 20-40%, intercepting literal duplicates. - Semantic cache (L2): embed the prompt, store it in a vector store. New requests embed first, retrieve top-k similar prompts, and if cosine similarity > 0.92 and the cached response is still fresh, return it directly. This catches "same question, different wording."
- Warm cache (L3, optional): pre-generate responses offline for high-frequency queries (product FAQ, doc summaries) and serve from a CDN — these never hit the gateway at all.
Hit Policy and Invalidation
- Cacheability check: requests with
temperature=0and notool_callsare cacheable; creative writing is not by default. - Tiered TTL: 24h for factual Q&A, 6h for code generation, 1h for time-sensitive content.
- Active invalidation: when a model version upgrades, clear the corresponding prefix; users see
x-cache: HIT/MISSin the response header. - Cache poisoning defense: prompts containing secrets (e.g.
key=xxx) are never cached.
Code Example
import hashlib, redis
r = redis.Redis()
from openai import OpenAI
emb_client = OpenAI()
def cache_key(messages, model):
return "v1:" + hashlib.sha256(
(str(messages) + model).encode()).hexdigest()
def get_or_call(messages, model, call_fn):
k = cache_key(messages, model)
if cached := r.get(k):
return cached
# semantic check: embed and search neighbors
emb = emb_client.embeddings.create(
model="text-embedding-3-small",
input=str(messages)).data[0].embedding
if near := vector_db.search(emb, threshold=0.92):
r.setex(k, 3600, near)
return near
resp = call_fn()
r.setex(k, 3600, resp)
vector_db.upsert(emb, resp)
return resp
Embedding Model Choice
Semantic cache hit rate depends on the embedding model. OpenAI text-embedding-3-small (1536 dims) hits well but has cost; local bge-small-zh is free but slightly less accurate. Key trap: once you change the embedding model, the entire vector store becomes invalid and must be re-indexed. So the gateway must support "dual indexing" — the new model builds a new store while the old one keeps serving until its hit rate reaches zero, then is taken offline.
Best Practices
- Measure hit rate: a rate below 10% usually means the cache key is too coarse or the TTL too short.
- Bound cache size: LRU with a 1GB ceiling to avoid blowing up Redis.
- Provide a bypass: for debugging, expose an
x-bypass-cache: trueheader that skips the cache. - Cross-user isolation: never share cache entries across
client_idboundaries to avoid sensitive data leaks.
Caching turns free quota into "unlimited" — as long as your question distribution has a long tail.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key