← Back to articles
Detailed Usage

NVIDIA NIM Inference Example

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Introduction

NIM endpoints are fully OpenAI Chat Completions-compatible, which means you can reuse the familiar openai Python SDK — just swap the base_url and api_key. This article covers four common patterns: single-turn, streaming, multi-turn, and structured JSON output.

架构图

flowchart TD A[OpenAI SDK] --> B[NIM endpoint] B --> C[Chat completion] B --> D[Streaming] B --> E[Multi-turn] B --> F[JSON mode] C --> G[Response] D --> G E --> G F --> G

Setup

pip install openai

Set environment variables (local or cloud — pick one):

# Local NIM container
export NIM_BASE_URL="http://localhost:8000/v1"
export NIM_API_KEY="local-no-key-needed"

# Cloud build.nvidia.com
export NIM_BASE_URL="https://integrate.api.nvidia.com/v1"
export NIM_API_KEY="nvapi-..."

Example 1: Single-turn Chat

import os
from openai import OpenAI

client = OpenAI(
    base_url=os.environ["NIM_BASE_URL"],
    api_key=os.environ["NIM_API_KEY"],
)

resp = client.chat.completions.create(
    model="meta/llama-3.1-8b-instruct",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Explain quantum entanglement in one sentence."},
    ],
    temperature=0.5,
    max_tokens=128,
)
print(resp.choices[0].message.content)
print("tokens:", resp.usage.total_tokens)

Example 2: Streaming

Streaming dramatically reduces time-to-first-token for long answers:

stream = client.chat.completions.create(
    model="meta/llama-3.1-8b-instruct",
    messages=[{"role": "user", "content": "Write a haiku about autumn."}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)
print()

Example 3: Multi-turn Conversation

history = [
    {"role": "system", "content": "You are a patient programming tutor."}
]
def chat(user_text):
    history.append({"role": "user", "content": user_text})
    r = client.chat.completions.create(
        model="meta/llama-3.1-8b-instruct",
        messages=history,
    )
    msg = r.choices[0].message
    history.append(msg)
    return msg.content

print(chat("What is the difference between list and tuple in Python?"))
print(chat("When should I prefer a tuple?"))

Example 4: Structured JSON Output

When you need machine-parseable JSON, declare response_format explicitly:

import json

resp = client.chat.completions.create(
    model="meta/llama-3.1-8b-instruct",
    messages=[
        {"role": "system", "content": "You are a data extraction assistant. Output JSON only."},
        {"role": "user", "content": "Extract name and title: 'Alice is a Senior Engineer at Acme Corp.'"},
    ],
    response_format={"type": "json_object"},
    max_tokens=128,
)
data = json.loads(resp.choices[0].message.content)
print(data)
# {'name': 'Alice', 'title': 'Senior Engineer', 'company': 'Acme Corp.'}

Performance Tuning

  • Enable streaming: stream=True brings time-to-first-token under 100ms.
  • Reuse the client: make OpenAI() a singleton to avoid rebuilding the connection pool per request.
  • Set timeouts: client = OpenAI(..., timeout=30, max_retries=2) to prevent long-tail hangs.
  • Batch requests: NIM supports /v1/batch (some models) for offline workloads.

Advanced: Tool Use and Batching

NIM also supports OpenAI-style function calling, ideal for agent orchestration:

tools = [{
    "type": "function",
    "function": {
        "name": "get_stock",
        "description": "Get stock price",
        "parameters": {
            "type": "object",
            "properties": {"symbol": {"type": "string"}},
            "required": ["symbol"],
        },
    },
}]
resp = client.chat.completions.create(
    model="meta/llama-3.1-8b-instruct",
    messages=[{"role":"user","content":"Check AAPL"}],
    tools=tools,
)
print(resp.choices[0].message.tool_calls[0].function)

For offline batch jobs, some NIM images also expose /v1/batch so you can submit thousands of prompts at once, billed by completion — about 50% cheaper than per-call.

Troubleshooting

  • model not found: Local NIM model names are baked into the image. Query GET /v1/models to see the actual ID.
  • Slow cloud responses: The first call may take 1-2 seconds for cold start; subsequent calls settle to hundreds of milliseconds.
  • JSON parse failures: Strongly constrain the format in the system prompt, e.g. "Output JSON only with fields name/title/company".

Master these four patterns and you have most business scenarios covered.

Best Practices

  • Prefer streaming: NIM streaming is ~30% faster than non-streaming overall, with better UX.
  • Use JSON mode for structured output: pass response_format={"type":"json_object"} to force JSON and avoid parse failures.
  • Tune temperature for the task: 0.7 default suits chat; 0.3 for code; 0.5 for summarization.
  • Leave 1024-token headroom in max_tokens: models often truncate long answers; leave headroom to avoid being cut off.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free AI Content Moderation API Guide 2026: Llama Guard 3 vs Perspective vs OpenAIFree OCR and Document Parsing API in PracticeIntegrating Free APIs into Your Local IDEConnecting Free Models to OpenCode in PracticeApplying for an OpenRouter API Key and Understanding Pricing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.