← Back to articles
Unified API Calling

LiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models — OpenAI Compatible + Smart Routing

LiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models

TL;DR: LiteLLM Proxy (litellm/proxy) collapses 100+ LLMs (OpenAI, Claude, Gemini, DeepSeek, Groq, Bedrock, Vertex) behind one OpenAI-compatible endpoint with load balancing, retries, budgets and observability — zero-change for Python teams.

Why LiteLLM?

  • 100+ models: one model name for OpenAI/Anthropic/Google/DeepSeek/Groq/Bedrock/Azure/Vertex.
  • OpenAI compat: POST /v1/chat/completions drop-in.
  • Smart routing: weight/latency/quota/cost, auto circuit break.
  • Budgets: per-key/team/model budgets, RPM/TPM limits.
  • Observability: Prometheus + Langfuse + logs.

Quick Start

pip install litellm
litellm proxy --config litellm_config.yaml --port 4000
from openai import OpenAI
client = OpenAI(base_url="http://localhost:4000/v1", api_key="sk-litellm")
client.chat.completions.create(model="deepseek-chat", messages=[{"role":"user","content":"hello"}])

Unified Calling

  • Load balance: routing_strategy: latency-based picks fastest.
  • Budgets: max_budget: 10 per team key.
  • Fallbacks: fallbacks: {gpt-4o-mini: [deepseek-chat, gemini-flash]} auto degrades.

vs OneAPI / Portkey

Solution Lang Models Best For
LiteLLM Python 100+ Python teams
OneAPI Go 10+ Self-host + monetize
Portkey Node/Go 250+ Enterprise

Best Practices

  • Weight free models higher, paid as fallback.
  • Standardize timeouts 30s/120s.
  • Prometheus + Grafana for p95/cost.

Summary

  • Zero-change Python: one base_url change.
  • Smart routing: latency/cost/quota.
  • Stack with OneAPI: upstream → OneAPI mymodel.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App — Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling — Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool — Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models — Cache + Guardrails + ObservabilityNextChat Complete Guide: Lightweight Web Unified Calling — One-Click Model Switch + Prompt Marketplace

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.