⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Introduction
Groq, Together AI, DeepSeek, NVIDIA NIM, and OpenRouter all implement the OpenAI Chat Completions protocol. That means a single openai SDK codebase — with only base_url and api_key swapped — can move seamlessly across all of them. This article provides a unified wrapper and demos streaming and async patterns.
架构图
Install
pip install openai python-dotenv
Unified Wrapper
import os
from dataclasses import dataclass
from openai import OpenAI
from dotenv import load_dotenv
load_dotenv()
@dataclass
class Provider:
name: str
base_url: str
api_key_env: str
model: str
PROVIDERS = {
"groq": Provider("Groq", "https://api.groq.com/openai/v1", "GROQ_API_KEY", "llama-3.3-70b-versatile"),
"together": Provider("Together", "https://api.together.xyz/v1", "TOGETHER_API_KEY", "meta-llama/Llama-3.3-70B-Instruct-Turbo"),
"deepseek": Provider("DeepSeek", "https://api.deepseek.com/v1", "DEEPSEEK_API_KEY", "deepseek-chat"),
"nim": Provider("NIM", "https://integrate.api.nvidia.com/v1", "NVIDIA_API_KEY", "meta/llama-3.1-8b-instruct"),
"openrouter": Provider("OpenRouter", "https://openrouter.ai/api/v1", "OPENROUTER_API_KEY","meta-llama/llama-3.3-70b-instruct:free"),
}
def client_for(name: str) -> OpenAI:
p = PROVIDERS[name]
return OpenAI(api_key=os.environ[p.api_key_env], base_url=p.base_url)
def chat(name: str, prompt: str, system: str = "You are a concise assistant.") -> str:
p = PROVIDERS[name]
resp = client_for(name).chat.completions.create(
model=p.model,
messages=[
{"role": "system", "content": system},
{"role": "user", "content": prompt},
],
temperature=0.3,
max_tokens=256,
)
return resp.choices[0].message.content
if __name__ == "__main__":
for name in PROVIDERS:
try:
print(f"[{name}] {chat(name, 'Explain an inverted index in one sentence.')}")
except Exception as e:
print(f"[{name}] ERROR: {e}")
Streaming Wrapper
def chat_stream(name: str, prompt: str):
p = PROVIDERS[name]
stream = client_for(name).chat.completions.create(
model=p.model,
messages=[{"role": "user", "content": prompt}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
yield delta
for piece in chat_stream("groq", "Write a poem about autumn."):
print(piece, end="", flush=True)
print()
Async Concurrency
For batch jobs, AsyncOpenAI lets you fan out across providers:
import asyncio
from openai import AsyncOpenAI
async def achat(name, prompt):
p = PROVIDERS[name]
client = AsyncOpenAI(api_key=os.environ[p.api_key_env], base_url=p.base_url)
r = await client.chat.completions.create(
model=p.model,
messages=[{"role":"user","content":prompt}],
max_tokens=64,
)
return name, r.choices[0].message.content
async def main():
results = await asyncio.gather(*[achat(n, "hi") for n in PROVIDERS])
for name, text in results:
print(f"[{name}] {text}")
asyncio.run(main())
Selection Strategy
- Lowest latency: Groq (LPU)
- Cheapest: OpenRouter
:freemodels - Best Chinese: DeepSeek
- On-prem: NIM container
- Multi-provider failover: OpenRouter's
fallbackfield
Failover and Degradation
In production, combine providers: use a stable paid provider (OpenAI/Anthropic) on the main path, with several free providers as backups that activate on failure. OpenRouter's fallback field gives zero-code failover; alternatively wrap calls in try/except for manual switching. Use asyncio.Semaphore to cap concurrency in async mode; for batch processing, push requests into a queue and have workers consume at each provider's rate limit. For cross-region workloads, overseas providers see higher latency variance at peak hours — reuse connections and route via a local proxy.
Streaming Wrapper
All providers support OpenAI-style streaming responses, so one generator wraps them all. The snippet below turns streamed tokens into a Python generator, so the business layer stays provider-agnostic:
def chat_stream(name: str, prompt: str):
p = PROVIDERS[name]
stream = client_for(name).chat.completions.create(
model=p.model,
messages=[{"role": "user", "content": prompt}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
yield delta
for piece in chat_stream("groq", "Write a poem about autumn."):
print(piece, end="", flush=True)
print()
This single generator can power SSE push, CLI typewriter effects, and more.
Troubleshooting
401: The env var for the provider is unset or misspelled. Checkapi_key_env.model not found: Model IDs differ across platforms. Trust each platform's/v1/modelsresponse.- Auto fallback: Wrap calls in try/except and switch to the next provider on failure.
- Token accounting differs: Providers vary on whether system prompts count toward cache. Normalize the baseline before comparing costs.
Master this pattern and switching models becomes a one-line change.
Best Practices
- Abstract a
make_client(provider)factory: returnOpenAI(base_url, api_key)uniformly; business code stays provider-agnostic. - Name env vars per provider:
GROQ_API_KEY,TOGETHER_API_KEY,DEEPSEEK_API_KEY, etc. — separate. base_urlmust end with/v1: missing the/v1returns 404; all compatibility layers follow OpenAI path conventions.- Normalize error codes: 429 means different things across providers (RPM vs TPM); the gateway must normalize.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key