← Back to articles
Free API Overview

NVIDIA NIM Free API — 83 Models (57 Chat + 26 Specialty)

NVIDIA NIM Free API Details

Verified on 2026-08-27 00:35 +08:00 (source: OneAPI production DB Channel table direct read + official API https://integrate.api.nvidia.com/v1/models no-auth verified 83) Channel slug:nvidia-nim | Detail URL:/free-llm-api-rankings/nvidia-nim | Base URL:https://integrate.api.nvidia.com/v1


1. Channel Overview

NVIDIA NIM (NVIDIA Inference Microservice) exposes 83 free models via integrate.api.nvidia.com/v1 — verified 2026-08-27 00:35 without auth (83 total, 33 nvidia/ official). The free tier covers chat (57), embedding (7), moderation/safety (5), translation (3) and vision (6), including Llama, Gemma, Qwen, Mistral, DeepSeek, Kimi, Phi and the full Nemotron family. 60 RPM, OpenAI-compatible /v1/chat/completions and /v1/models.


2. Verified Supported Models

Scope: models_total = 83 (full Channel.models direct read + official /v1/models 83 cross-verified); chat = 57; embedding = 7; moderation = 5; translation = 3; vision = 6

Chat Models (57)

  • 01-ai/yi-large
  • ai21labs/jamba-1.5-large-instruct
  • aisingapore/sea-lion-7b-instruct
  • bigcode/starcoder2-15b
  • databricks/dbrx-instruct
  • deepseek-ai/deepseek-coder-6.7b-instruct
  • deepseek-ai/deepseek-v4-flash-0731
  • deepseek-ai/deepseek-v4-pro-0813
  • google/codegemma-1.1-7b
  • google/codegemma-7b
  • google/diffusiongemma-26b-a4b-it
  • google/gemma-2b
  • google/gemma-3-12b-it
  • google/gemma-3-4b-it
  • google/gemma-4-31b-it
  • google/recurrentgemma-2b
  • ibm/granite-3.0-3b-a800m-instruct
  • ibm/granite-3.0-8b-instruct
  • ibm/granite-34b-code-instruct
  • ibm/granite-8b-code-instruct
  • meta/codellama-70b
  • meta/llama2-70b
  • meta/muse-glimmer-30b
  • microsoft/phi-3.5-moe-instruct
  • minimaxai/minimax-m3
  • mistralai/codestral-22b-instruct-v0.1
  • mistralai/mistral-7b-instruct-v0.3
  • mistralai/mistral-large
  • mistralai/mistral-large-2-instruct
  • mistralai/mistral-nemotron
  • mistralai/mixtral-8x22b-v0.1
  • moonshotai/kimi-k2.6
  • moonshotai/kimi-k3
  • nv-mistralai/mistral-nemo-12b-instruct
  • nvidia/cosmos-reason2-8b
  • nvidia/ising-calibration-1.5-31b
  • nvidia/llama-3.1-nemotron-51b-instruct
  • nvidia/llama-3.1-nemotron-70b-instruct
  • nvidia/llama-3.1-nemotron-ultra-253b-v1
  • nvidia/llama3-chatqa-1.5-70b
  • nvidia/mistral-nemo-minitron-8b-8k-instruct
  • nvidia/nemotron-3-nano-30b-a3b
  • nvidia/nemotron-3-nano-omni-30b-a3b-reasoning
  • nvidia/nemotron-3-super-120b-a12b
  • nvidia/nemotron-3-ultra-550b-a55b
  • nvidia/nemotron-3.5-lightning-30b-a3b
  • nvidia/nemotron-4-340b-instruct
  • nvidia/nemotron-4-340b-reward
  • nvidia/nemotron-nano-3-30b-a3b
  • openai/gpt-oss-120b
  • openai/gpt-oss-20b
  • poolside/laguna-xs-2.1
  • writer/palmyra-creative-122b
  • writer/palmyra-fin-70b-32k
  • writer/palmyra-med-70b
  • writer/palmyra-med-70b-32k
  • zyphra/zamba2-7b-instruct

Embedding Models (7)

  • nvidia/embed-qa-4
  • nvidia/llama-3.2-nemoretriever-1b-vlm-embed-v1
  • nvidia/llama-3.2-nv-embedqa-1b-v1
  • nvidia/llama-nemotron-embed-vl-1b-v2
  • nvidia/nemotron-3-embed-1b
  • nvidia/nv-embedqa-mistral-7b-v2
  • snowflake/arctic-embed-l

Moderation/Safety Models (5)

  • meta/llama-guard-4-12b
  • nvidia/llama-3.1-nemoguard-8b-content-safety
  • nvidia/llama-3.1-nemoguard-8b-topic-control
  • nvidia/llama-3.1-nemotron-safety-guard-8b-v3
  • nvidia/nemotron-3.5-content-safety

Translation Models (3)

  • nvidia/riva-translate-4b-instruct
  • nvidia/riva-translate-4b-instruct-v1.1
  • nvidia/riva-translate-4b-instruct-v2

Vision Models (6)

  • meta/llama-3.2-11b-vision-instruct
  • meta/llama-3.2-90b-vision-instruct
  • microsoft/phi-3-vision-128k-instruct
  • nvidia/cosmos-reason2-8b
  • nvidia/neva-22b
  • nvidia/vila

3. Registration

  1. Visit https://build.nvidia.com and sign in with NVIDIA account.
  2. Generate API key at https://integrate.api.nvidia.com (header Authorization: Bearer nvapi-...).
  3. Free tier is available without credit card; rate limit 60 RPM.

4. Usage

Base URL: https://integrate.api.nvidia.com/v1

OpenAI-compatible. List models: GET /v1/models (no auth required, returns 83). Chat: POST /v1/chat/completions with model and messages. Example: model: "openai/gpt-oss-120b" or "nvidia/nemotron-3-nano-30b-a3b".


5. Verified Metrics

实测于 2026-08-27 00:35 +08:00,官方 /v1/models 无鉴权 200,返回 83。

Metric Value Notes
HTTP 状态码 200 GET /v1/models 无鉴权
模型总数 83 含 33 个 nvidia/ 官方
对话模型 57 含 Llama/Qwen/Mistral/DeepSeek 等
首 token 延迟 待填 需 /chat/completions 实测
RPM 限制 60 Channel 表 maxRpm
免费层 ✅ 无需信用卡,build.nvidia.com

6. Pros & Cons

Pros:

  • 83 models, largest free pool (57 chat + 6 vision + 7 embedding)
  • 33 nvidia/ official Nemotron models
  • OpenAI-compatible, no auth for /v1/models
  • GPU-accelerated NIM

Cons:

  • Some models EOL (e.g. meta/llama-3.1-70b-instruct 410 on 2026-08-26)
  • Translation/vision models are specialized, not general chat

本页由 IT 于 2026-08-27 00:35 四重复验生成,数据以 oneapi 生产 DB Channel 表直读 + 官方 API 实测为准。


🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →

Start on APIShare in three steps

Ready to try the options above? Three steps get you running:

  1. Create an account - open the APIShare free API registration page. An email address is all you need; no credit card required.
  2. Browse the free API catalog - head to the complete free API list and filter by text, image, audio, embedding, or multimodal. Each entry shows its free quota, rate limit, and availability status.
  3. Grab a key and integrate - generate an API key in your dashboard and paste it into your application. Every plan includes actively-updated APIs gateways covering every provider mentioned in this guide.

Already have an account? Use the APIShare login page, or visit the APIShare homepage for a full platform overview. Registration is free, and you can stop at any time.

Every outbound link in this guide carries a UTM parameter (utm_source=apishare_devto&utm_medium=referral&utm_campaign=free_api_article) for clean campaign attribution.


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free API Cost and Quota Control in Practice: 429 Backoff, RPM Budgets, and Multi-Model Fallback2026 Free OneAPI Unified Gateway: Connect 100+ LLM APIs at Zero Cost in One GuideRun a 550B-Parameter Model for Free: 2026 Nemotron 3 Ultra Complete Guide (OpenRouter Free Tier Tested)2026 Free Embedding Vector Model API Panorama: BGE-M3 / Voyage / Nomic / Google / Azure and 6 Options Tested (September Update)Free Function Calling / Tool Use API Tutorial: DeepSeek / Gemini / Qwen — Zero-Cost Agent Tooling (2026-09-16 Verified)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.