⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Introduction
NVIDIA NIM (NVIDIA Inference Microservices) is a containerized inference service that bundles optimizations like TensorRT-LLM and Triton behind an OpenAI-compatible API. You can either run NIM locally with Docker or call hosted endpoints on build.nvidia.com. This article demonstrates both and gives guidance on choosing between them.
架构图
Option 1: Local Container Deployment
Prerequisites: an NVIDIA GPU (>= 16 GB VRAM) with driver >= 535, nvidia-container-toolkit, and an NGC account.
# 1. Log in to NGC and pull the image
docker login nvcr.io -u '$oauthtoken' -p "$(cat ngc-key.txt)"
# 2. Pull the Llama-3.1-8B NIM image
docker pull nvcr.io/nim/meta/llama-3.1-8b-instruct:1.0.0
# 3. Start the container (single GPU)
docker run -d --name nim-llama \
--runtime=nvidia --gpus all \
-e NGC_API_KEY=$(cat ngc-key.txt) \
-v $HOME/.cache/nim:/opt/nim/.cache \
-p 8000:8000 \
--shm-size=16g \
nvcr.io/nim/meta/llama-3.1-8b-instruct:1.0.0
The first launch downloads the model weights (~16 GB). When ready, visit http://localhost:8000/docs for the Swagger UI. /v1/models lists available models, and /v1/chat/completions is the OpenAI-compatible endpoint.
Option 2: Call Cloud Endpoints
Skip the GPU and use the hosted endpoints on build.nvidia.com. Register at https://build.nvidia.com, generate an NGC API key, then call:
curl https://integrate.api.nvidia.com/v1/chat/completions \
-H "Authorization: Bearer $NVIDIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "meta/llama-3.1-8b-instruct",
"messages": [{"role":"user","content":"hi"}],
"max_tokens": 128
}'
New accounts receive 1000 free calls — perfect for quick prototyping. The cloud also hosts specialized models like nvidia/llama-3.1-nemotron-70b-instruct (NVIDIA's aligned variant) and deepseek-ai/deepseek-r1 (a reasoning model).
Choosing Between Modes
| Dimension | Local NIM | Cloud NIM |
|---|---|---|
| Setup cost | High (needs A10/A100) | None |
| Latency | Low (local GPU) | Medium (over the internet) |
| Privacy | Fully on-prem | Data leaves the network |
| Best for | Internal/compliance workloads | Prototyping / low traffic |
Performance Tuning
- Enable TensorRT-LLM: enabled by default in NIM images. The first launch compiles the engine (~5 min), after which inference is 2-5x faster.
- Tune batch size: set
NIM_MAX_BATCH_SIZE=128in container env to boost throughput. - Multi-GPU inference: use
--gpus allplusNIM_TENSOR_PARALLEL_SIZE=2to shard the model across two GPUs.
Troubleshooting
CUDA out of memory: The 8B model needs at least 16 GB of VRAM. Retry with--shm-size=16gor switch to a 4B model.- 401 on image pull: Verify your NGC API key is valid and your account is enrolled in the NIM early access program.
- Other models: Browse all NIM images at https://build.nvidia.com and swap
nvcr.io/nim/<vendor>/<model>accordingly. - Slow cold start: The first start downloads weights and compiles the engine. Use
--restart unless-stoppedin production to keep the container warm.
Once deployed, see "NVIDIA NIM Inference Example" to call it via the OpenAI SDK.
Deployment Comparison
| Path | Strength | Limitation |
|---|---|---|
| Cloud build.nvidia.com | Zero ops, 1000 credits | Shared QPS, variable latency |
| Local Docker NIM | Dedicated GPU, low latency | Needs GPU card (A10/L4 minimum) |
Best Practices
- Start with cloud: burn through the 1000 credits to validate the flow before deciding on local.
- Pick L4 over A10 locally: L4 has better price/perf; a single card runs 7B-70B quantized comfortably.
- GPU driver prerequisite: install nvidia-docker runtime, otherwise the container cannot see the GPU.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key