⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Introduction
NIM endpoints are fully OpenAI Chat Completions-compatible, which means you can reuse the familiar openai Python SDK — just swap the base_url and api_key. This article covers four common patterns: single-turn, streaming, multi-turn, and structured JSON output.
架构图
Setup
pip install openai
Set environment variables (local or cloud — pick one):
# Local NIM container
export NIM_BASE_URL="http://localhost:8000/v1"
export NIM_API_KEY="local-no-key-needed"
# Cloud build.nvidia.com
export NIM_BASE_URL="https://integrate.api.nvidia.com/v1"
export NIM_API_KEY="nvapi-..."
Example 1: Single-turn Chat
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ["NIM_BASE_URL"],
api_key=os.environ["NIM_API_KEY"],
)
resp = client.chat.completions.create(
model="meta/llama-3.1-8b-instruct",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain quantum entanglement in one sentence."},
],
temperature=0.5,
max_tokens=128,
)
print(resp.choices[0].message.content)
print("tokens:", resp.usage.total_tokens)
Example 2: Streaming
Streaming dramatically reduces time-to-first-token for long answers:
stream = client.chat.completions.create(
model="meta/llama-3.1-8b-instruct",
messages=[{"role": "user", "content": "Write a haiku about autumn."}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
print()
Example 3: Multi-turn Conversation
history = [
{"role": "system", "content": "You are a patient programming tutor."}
]
def chat(user_text):
history.append({"role": "user", "content": user_text})
r = client.chat.completions.create(
model="meta/llama-3.1-8b-instruct",
messages=history,
)
msg = r.choices[0].message
history.append(msg)
return msg.content
print(chat("What is the difference between list and tuple in Python?"))
print(chat("When should I prefer a tuple?"))
Example 4: Structured JSON Output
When you need machine-parseable JSON, declare response_format explicitly:
import json
resp = client.chat.completions.create(
model="meta/llama-3.1-8b-instruct",
messages=[
{"role": "system", "content": "You are a data extraction assistant. Output JSON only."},
{"role": "user", "content": "Extract name and title: 'Alice is a Senior Engineer at Acme Corp.'"},
],
response_format={"type": "json_object"},
max_tokens=128,
)
data = json.loads(resp.choices[0].message.content)
print(data)
# {'name': 'Alice', 'title': 'Senior Engineer', 'company': 'Acme Corp.'}
Performance Tuning
- Enable streaming:
stream=Truebrings time-to-first-token under 100ms. - Reuse the client: make
OpenAI()a singleton to avoid rebuilding the connection pool per request. - Set timeouts:
client = OpenAI(..., timeout=30, max_retries=2)to prevent long-tail hangs. - Batch requests: NIM supports
/v1/batch(some models) for offline workloads.
Advanced: Tool Use and Batching
NIM also supports OpenAI-style function calling, ideal for agent orchestration:
tools = [{
"type": "function",
"function": {
"name": "get_stock",
"description": "Get stock price",
"parameters": {
"type": "object",
"properties": {"symbol": {"type": "string"}},
"required": ["symbol"],
},
},
}]
resp = client.chat.completions.create(
model="meta/llama-3.1-8b-instruct",
messages=[{"role":"user","content":"Check AAPL"}],
tools=tools,
)
print(resp.choices[0].message.tool_calls[0].function)
For offline batch jobs, some NIM images also expose /v1/batch so you can submit thousands of prompts at once, billed by completion — about 50% cheaper than per-call.
Troubleshooting
model not found: Local NIM model names are baked into the image. QueryGET /v1/modelsto see the actual ID.- Slow cloud responses: The first call may take 1-2 seconds for cold start; subsequent calls settle to hundreds of milliseconds.
- JSON parse failures: Strongly constrain the format in the system prompt, e.g. "Output JSON only with fields name/title/company".
Master these four patterns and you have most business scenarios covered.
Best Practices
- Prefer streaming: NIM streaming is ~30% faster than non-streaming overall, with better UX.
- Use JSON mode for structured output: pass
response_format={"type":"json_object"}to force JSON and avoid parse failures. - Tune temperature for the task: 0.7 default suits chat; 0.3 for code; 0.5 for summarization.
- Leave 1024-token headroom in max_tokens: models often truncate long answers; leave headroom to avoid being cut off.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key