โ† Back to articles
Unified API Calling

Privately Deployed Unified Gateway

โš ๏ธ Pending Update ยท 2026-08-29 Verification ยท Content may be outdated, please refer to official docs Updated: 2026-08-29 ยท Status: Pending Verification

Background

In healthcare, finance, and government scenarios, data cannot leave the intranet; offline or edge devices (satcom, factory floors) have no public-internet access at all. Deploying open-source models locally via Ollama or vLLM, then registering them as "internal providers" alongside cloud free models in the unified gateway, satisfies both compliance and offline degradation requirements.

Core Architecture

  • Hybrid provider pool: local ollama/llama3 and cloud openrouter/deepseek both sit in the routing table. The gateway routes by data-sensitivity tag โ€” sensitive requests are forced to the local pool.
  • Local provider onboarding: Ollama exposes http://localhost:11434/v1 and vLLM exposes http://localhost:8000/v1; both are natively OpenAI-compatible, so onboarding needs zero code changes.
  • Offline degradation chain: cloud-primary โ†’ local-fallback โ†’ cache-only. When the internet is up, prefer cloud free; when offline, switch to local; when local cannot run either, return cached responses.
  • Resource governance: the local GPU is scarce โ€” cap concurrency (e.g. 2 concurrent) and spill to cloud after a 30s queue.

Code Example

PROVIDERS = {
    "local": {"base": "http://localhost:11434/v1", "key": "ollama",
              "sensitive_only": True, "max_concurrent": 2},
    "cloud": {"base": "https://openrouter.ai/api/v1", "key": "sk-or-..."},
}

sem = asyncio.Semaphore(PROVIDERS["local"]["max_concurrent"])

async def route(req, is_sensitive):
    if is_sensitive:
        async with sem:
            return await call(PROVIDERS["local"], req)
    try:
        return await call(PROVIDERS["cloud"], req)
    except (NetworkError, TimeoutError):
        # offline fallback to local
        async with sem:
            return await call(PROVIDERS["local"], req)

Model Version Pinning

Locally deployed models must be explicitly version-pinned. Ollama pulls latest by default, but llama3:latest last week and this week may be different fine-tunes, causing production drift. Production must use full tags like llama3:8b-instruct-q5_k_M. Maintain a "model manifest" YAML recording each local model's exact version, quantization method, and context length; changes go through PR review โ€” never let "it worked yesterday, the model changed itself today" happen.

Best Practices

  • Local model selection: use a 7B model for inference tasks (runs on a single local GPU) and a 70B model (multiple GPUs or quantized) for complex tasks; route by model_size.
  • Data redaction: even when traffic stays local, do not log plaintext prompts โ€” log leakage remains a risk.
  • Resource monitoring: track local GPU utilization, VRAM, and queue depth in the gateway metrics dashboard so the local provider does not become an invisible bottleneck.
  • Canary: a newly deployed local provider first absorbs 1% of traffic, with success rate and latency under observation, before ramping up.
  • Cold-start warm-up: a local large model needs 30s+ on first load; the gateway pings it on startup to pre-warm so the first real request is not stuck behind a cold start.

Private deployment is not "local only" โ€” it is "sensitive local, ordinary cloud, gateway as conductor."

๐Ÿš€ Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key โ€” one key, 100+ models, free models at zero cost.

๐Ÿ‘‰ Register on Apishare.cc โ†’ Get your unified API Key

๐Ÿ“Š Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings โ†’


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App โ€” Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling โ€” Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool โ€” Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models โ€” Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models โ€” OpenAI Compatible + Smart Routing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide โ€” sign up and get bonus credits.