← Back to articles
Detailed Usage

Calling Free APIs via the OpenAI SDK Compatibility Layer

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Introduction

Groq, Together AI, DeepSeek, NVIDIA NIM, and OpenRouter all implement the OpenAI Chat Completions protocol. That means a single openai SDK codebase — with only base_url and api_key swapped — can move seamlessly across all of them. This article provides a unified wrapper and demos streaming and async patterns.

架构图

flowchart TD A[OpenAI SDK codebase] --> B{Switch base_url + api_key} B -->|groq.com| C[Groq LPU] B -->|together.xyz| D[Together AI] B -->|deepseek.com| E[DeepSeek] B -->|integrate.api.nvidia.com| F[NVIDIA NIM] C --> G[Same SDK shape] D --> G E --> G F --> G

Install

pip install openai python-dotenv

Unified Wrapper

import os
from dataclasses import dataclass
from openai import OpenAI
from dotenv import load_dotenv

load_dotenv()

@dataclass
class Provider:
    name: str
    base_url: str
    api_key_env: str
    model: str

PROVIDERS = {
    "groq":       Provider("Groq",       "https://api.groq.com/openai/v1",       "GROQ_API_KEY",      "llama-3.3-70b-versatile"),
    "together":   Provider("Together",   "https://api.together.xyz/v1",          "TOGETHER_API_KEY",  "meta-llama/Llama-3.3-70B-Instruct-Turbo"),
    "deepseek":   Provider("DeepSeek",   "https://api.deepseek.com/v1",          "DEEPSEEK_API_KEY",  "deepseek-chat"),
    "nim":        Provider("NIM",        "https://integrate.api.nvidia.com/v1",  "NVIDIA_API_KEY",    "meta/llama-3.1-8b-instruct"),
    "openrouter": Provider("OpenRouter", "https://openrouter.ai/api/v1",         "OPENROUTER_API_KEY","meta-llama/llama-3.3-70b-instruct:free"),
}

def client_for(name: str) -> OpenAI:
    p = PROVIDERS[name]
    return OpenAI(api_key=os.environ[p.api_key_env], base_url=p.base_url)

def chat(name: str, prompt: str, system: str = "You are a concise assistant.") -> str:
    p = PROVIDERS[name]
    resp = client_for(name).chat.completions.create(
        model=p.model,
        messages=[
            {"role": "system", "content": system},
            {"role": "user", "content": prompt},
        ],
        temperature=0.3,
        max_tokens=256,
    )
    return resp.choices[0].message.content

if __name__ == "__main__":
    for name in PROVIDERS:
        try:
            print(f"[{name}] {chat(name, 'Explain an inverted index in one sentence.')}")
        except Exception as e:
            print(f"[{name}] ERROR: {e}")

Streaming Wrapper

def chat_stream(name: str, prompt: str):
    p = PROVIDERS[name]
    stream = client_for(name).chat.completions.create(
        model=p.model,
        messages=[{"role": "user", "content": prompt}],
        stream=True,
    )
    for chunk in stream:
        delta = chunk.choices[0].delta.content
        if delta:
            yield delta

for piece in chat_stream("groq", "Write a poem about autumn."):
    print(piece, end="", flush=True)
print()

Async Concurrency

For batch jobs, AsyncOpenAI lets you fan out across providers:

import asyncio
from openai import AsyncOpenAI

async def achat(name, prompt):
    p = PROVIDERS[name]
    client = AsyncOpenAI(api_key=os.environ[p.api_key_env], base_url=p.base_url)
    r = await client.chat.completions.create(
        model=p.model,
        messages=[{"role":"user","content":prompt}],
        max_tokens=64,
    )
    return name, r.choices[0].message.content

async def main():
    results = await asyncio.gather(*[achat(n, "hi") for n in PROVIDERS])
    for name, text in results:
        print(f"[{name}] {text}")

asyncio.run(main())

Selection Strategy

  • Lowest latency: Groq (LPU)
  • Cheapest: OpenRouter :free models
  • Best Chinese: DeepSeek
  • On-prem: NIM container
  • Multi-provider failover: OpenRouter's fallback field

Failover and Degradation

In production, combine providers: use a stable paid provider (OpenAI/Anthropic) on the main path, with several free providers as backups that activate on failure. OpenRouter's fallback field gives zero-code failover; alternatively wrap calls in try/except for manual switching. Use asyncio.Semaphore to cap concurrency in async mode; for batch processing, push requests into a queue and have workers consume at each provider's rate limit. For cross-region workloads, overseas providers see higher latency variance at peak hours — reuse connections and route via a local proxy.

Streaming Wrapper

All providers support OpenAI-style streaming responses, so one generator wraps them all. The snippet below turns streamed tokens into a Python generator, so the business layer stays provider-agnostic:

def chat_stream(name: str, prompt: str):
    p = PROVIDERS[name]
    stream = client_for(name).chat.completions.create(
        model=p.model,
        messages=[{"role": "user", "content": prompt}],
        stream=True,
    )
    for chunk in stream:
        delta = chunk.choices[0].delta.content
        if delta:
            yield delta

for piece in chat_stream("groq", "Write a poem about autumn."):
    print(piece, end="", flush=True)
print()

This single generator can power SSE push, CLI typewriter effects, and more.

Troubleshooting

  • 401: The env var for the provider is unset or misspelled. Check api_key_env.
  • model not found: Model IDs differ across platforms. Trust each platform's /v1/models response.
  • Auto fallback: Wrap calls in try/except and switch to the next provider on failure.
  • Token accounting differs: Providers vary on whether system prompts count toward cache. Normalize the baseline before comparing costs.

Master this pattern and switching models becomes a one-line change.

Best Practices

  • Abstract a make_client(provider) factory: return OpenAI(base_url, api_key) uniformly; business code stays provider-agnostic.
  • Name env vars per provider: GROQ_API_KEY, TOGETHER_API_KEY, DEEPSEEK_API_KEY, etc. — separate.
  • base_url must end with /v1: missing the /v1 returns 404; all compatibility layers follow OpenAI path conventions.
  • Normalize error codes: 429 means different things across providers (RPM vs TPM); the gateway must normalize.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free AI Content Moderation API Guide 2026: Llama Guard 3 vs Perspective vs OpenAIFree OCR and Document Parsing API in PracticeIntegrating Free APIs into Your Local IDEConnecting Free Models to OpenCode in PracticeApplying for an OpenRouter API Key and Understanding Pricing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.