← Back to articles
Detailed Usage

NVIDIA NIM Local and Cloud Endpoint Deployment

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Introduction

NVIDIA NIM (NVIDIA Inference Microservices) is a containerized inference service that bundles optimizations like TensorRT-LLM and Triton behind an OpenAI-compatible API. You can either run NIM locally with Docker or call hosted endpoints on build.nvidia.com. This article demonstrates both and gives guidance on choosing between them.

架构图

flowchart LR A[NVIDIA NIM] --> B[Cloud: build.nvidia.com] A --> C[Local: Docker NIM container] B --> B1[Free 1000 credits] C --> C1[Self-hosted GPU box] B1 --> D[OpenAI-compatible endpoint] C1 --> D

Option 1: Local Container Deployment

Prerequisites: an NVIDIA GPU (>= 16 GB VRAM) with driver >= 535, nvidia-container-toolkit, and an NGC account.

# 1. Log in to NGC and pull the image
docker login nvcr.io -u '$oauthtoken' -p "$(cat ngc-key.txt)"

# 2. Pull the Llama-3.1-8B NIM image
docker pull nvcr.io/nim/meta/llama-3.1-8b-instruct:1.0.0

# 3. Start the container (single GPU)
docker run -d --name nim-llama \
  --runtime=nvidia --gpus all \
  -e NGC_API_KEY=$(cat ngc-key.txt) \
  -v $HOME/.cache/nim:/opt/nim/.cache \
  -p 8000:8000 \
  --shm-size=16g \
  nvcr.io/nim/meta/llama-3.1-8b-instruct:1.0.0

The first launch downloads the model weights (~16 GB). When ready, visit http://localhost:8000/docs for the Swagger UI. /v1/models lists available models, and /v1/chat/completions is the OpenAI-compatible endpoint.

Option 2: Call Cloud Endpoints

Skip the GPU and use the hosted endpoints on build.nvidia.com. Register at https://build.nvidia.com, generate an NGC API key, then call:

curl https://integrate.api.nvidia.com/v1/chat/completions \
  -H "Authorization: Bearer $NVIDIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta/llama-3.1-8b-instruct",
    "messages": [{"role":"user","content":"hi"}],
    "max_tokens": 128
  }'

New accounts receive 1000 free calls — perfect for quick prototyping. The cloud also hosts specialized models like nvidia/llama-3.1-nemotron-70b-instruct (NVIDIA's aligned variant) and deepseek-ai/deepseek-r1 (a reasoning model).

Choosing Between Modes

Dimension Local NIM Cloud NIM
Setup cost High (needs A10/A100) None
Latency Low (local GPU) Medium (over the internet)
Privacy Fully on-prem Data leaves the network
Best for Internal/compliance workloads Prototyping / low traffic

Performance Tuning

  • Enable TensorRT-LLM: enabled by default in NIM images. The first launch compiles the engine (~5 min), after which inference is 2-5x faster.
  • Tune batch size: set NIM_MAX_BATCH_SIZE=128 in container env to boost throughput.
  • Multi-GPU inference: use --gpus all plus NIM_TENSOR_PARALLEL_SIZE=2 to shard the model across two GPUs.

Troubleshooting

  • CUDA out of memory: The 8B model needs at least 16 GB of VRAM. Retry with --shm-size=16g or switch to a 4B model.
  • 401 on image pull: Verify your NGC API key is valid and your account is enrolled in the NIM early access program.
  • Other models: Browse all NIM images at https://build.nvidia.com and swap nvcr.io/nim/<vendor>/<model> accordingly.
  • Slow cold start: The first start downloads weights and compiles the engine. Use --restart unless-stopped in production to keep the container warm.

Once deployed, see "NVIDIA NIM Inference Example" to call it via the OpenAI SDK.

Deployment Comparison

Path Strength Limitation
Cloud build.nvidia.com Zero ops, 1000 credits Shared QPS, variable latency
Local Docker NIM Dedicated GPU, low latency Needs GPU card (A10/L4 minimum)

Best Practices

  • Start with cloud: burn through the 1000 credits to validate the flow before deciding on local.
  • Pick L4 over A10 locally: L4 has better price/perf; a single card runs 7B-70B quantized comfortably.
  • GPU driver prerequisite: install nvidia-docker runtime, otherwise the container cannot see the GPU.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free AI Content Moderation API Guide 2026: Llama Guard 3 vs Perspective vs OpenAIFree OCR and Document Parsing API in PracticeIntegrating Free APIs into Your Local IDEConnecting Free Models to OpenCode in PracticeApplying for an OpenRouter API Key and Understanding Pricing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.