← Back to articles
Unified API Calling

Unified Billing and Quota Management

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Background

Provider billing units are wildly inconsistent: OpenAI bills by token, Gemini by character, Replicate by second, and Cohere Trial by monthly request count. Free tiers pile on multi-dimensional limits — RPM, RPD, TPM, IPM. To produce a single "how much did the user spend" view, the gateway must normalize metering after each upstream response and write it back to a unified ledger.

Core Design

  • Unified unit: fold every upstream cost into "normalized tokens," where 1 credit = $1 equivalent. Free models record zero cost but still count call volume.
  • Three-layer ledger: user_quota (monthly total per user) → client_quota (daily per client_id) → provider_quota (free pool per provider). Layers decrement independently; the tightest layer triggers degradation first.
  • Free-first policy: routing sorts providers by remaining free quota descending; only when the free pool is empty does it switch to paid — this is "save every cent" as an executable policy.
  • Overage degradation: on hitting a layer cap, the gateway degrades the request to a cheaper model (e.g. 70B → 8B) instead of rejecting it.
  • Reconciliation: a nightly job pulls each provider's official billing API and compares to gateway metering; >5% deviation triggers an alert.

Code Example

class QuotaLedger:
    def __init__(self, redis):
        self.r = redis

    def consume(self, client_id, provider, tokens, is_free):
        # decrement three layers atomically
        pipe = self.r.pipeline()
        pipe.hincrby(f"u:{client_id}", "spent", tokens)
        pipe.hincrby(f"c:{client_id}:day", "spent", tokens)
        if is_free:
            pipe.hincrby(f"p:{provider}:free", "spent", tokens)
        pipe.execute()

    def route(self, client_id, candidates):
        # free first, sort by remaining free quota desc
        free = [(p, self.r.get(f"p:{p}:free:remain") or 0) for p in candidates]
        free.sort(key=lambda x: -int(x[1]))
        for p, _ in free:
            if self._can_use(p, client_id):
                return p, True
        return candidates[0], False  # fall back to paid

Real-Time vs Batch Metering

Billing has two granularities: real-time metering (write back per response) and batch metering (flush every 5 minutes). Real-time metering lets routing see "this provider's quota is nearly exhausted" immediately, but a Redis write per response has performance cost; batch metering has lower write pressure but routing cannot see the just-burned quota for 5 minutes. Compromise: accumulate in an in-memory counter, flush to Redis every 50 calls or 5 seconds, balancing latency with write pressure.

Best Practices

  • Budget guardrails: define a soft cap (degrade on overage) and a hard cap (reject on overage); the soft cap sits below the hard cap.
  • Cold-start credits: give new users 1000 normalized tokens up front, then prompt for a paid plan when exhausted — never just reject.
  • Pre-debit streaming: deduct incrementally as a streaming response emits tokens, not at the end, to prevent mid-stream overruns.
  • Automated reconciliation: a nightly job pulls each provider's billing API and reconciles against gateway metering; >5% deviation alerts.

Billing is not a back-office report — it is a real-time input to routing.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App — Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling — Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool — Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models — Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models — OpenAI Compatible + Smart Routing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.