← Back to articles
Free API Overview

Groq Free Ultra-Fast Inference API: 13 Free Models Tested

Introduction

Groq is famous for its custom LPU chip — a single card outputs hundreds of tokens per second, making it one of the lowest-latency inference clouds in 2026. This guide is based on APIShare's channel-pool verification (https://api.groq.com/openai/v1) on 2026-08-22: 13 free models were catalogued from the live endpoint, 11 of them verified working — 6 conversational models on the main ranking and 5 specialized-capability models listed separately, plus 2 Orpheus TTS models pending Groq's model-terms acceptance. Everything runs through a single OpenAI-compatible endpoint.

Main Ranking: Chat Models (6 · verified via chat-completions)

Model ID Context Max Output Features Best For
openai/gpt-oss-120b 128K 64K Reasoning / Tools / Structured Outputs Complex reasoning, agent brains
openai/gpt-oss-20b 128K 64K Reasoning / Tools Best value for high-volume calls
qwen/qwen3.6-27b 128K 16K Image input / Tools / Reasoning Multimodal, bilingual EN/ZH
groq/compound 128K 8K Agentic system (auto search & synthesis) Answers that need lookups
groq/compound-mini 128K 8K Lightweight agentic Low-cost automation pipelines
allam-2-7b 4K 4K json_mode Arabic-language scenarios

⚠️ Reasoning note: gpt-oss models and qwen3.6 generate a chain of thought first, which consumes output tokens. Set max_tokens ≥ 256 or you'll get thinking-only empty answers.

Specialized Models (7 · excluded from main chat ranking)

Model ID Capability Notes
whisper-large-v3-turbo Speech-to-text Speed-first, recommended default
whisper-large-v3 Speech-to-text Accuracy-first
canopylabs/orpheus-v1-english Text-to-speech English TTS · coming soon*
canopylabs/orpheus-arabic-saudi Text-to-speech Arabic TTS · coming soon*
openai/gpt-oss-safeguard-20b Safety classifier Output-side moderation
meta-llama/llama-prompt-guard-2-86m Input moderation Returns risk probability
meta-llama/llama-prompt-guard-2-22m Input moderation Lighter probability scorer

* Coming soon: requires a Groq org admin to accept the model terms in the console — calls currently return model_terms_required. The other 11 models are verified working.

Quick Start

from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key="gsk_...")
resp = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Explain why LPU is fast in one sentence"}],
    max_tokens=512)
print(resp.choices[0].message.content)

Rate Limits & Best Practices

  • Typical free-tier limits: ~30 RPM / 14,400 RPD. Add token-bucket backoff on 429s.
  • Free-tier data may be used for service improvement — avoid sensitive workloads.
  • Migration note: legacy llama-3.x models (llama-3.1-8b-instant, llama-3.3-70b-versatile, etc.) are no longer in the current channel pool. Migrate to openai/gpt-oss-20b or groq/compound-mini.
  • Fetch the live catalog anytime via GET /openai/v1/models; the APIShare unified endpoint already mirrors this catalog — 11 models ready out of the box, with the 2 Orpheus TTS models opening once activated.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →

Start on APIShare in three steps

Ready to try the options above? Three steps get you running:

  1. Create an account - open the APIShare free API registration page. An email address is all you need; no credit card required.
  2. Browse the free API catalog - head to the complete free API list and filter by text, image, audio, embedding, or multimodal. Each entry shows its free quota, rate limit, and availability status.
  3. Grab a key and integrate - generate an API key in your dashboard and paste it into your application. Every plan includes actively-updated APIs gateways covering every provider mentioned in this guide.

Already have an account? Use the APIShare login page, or visit the APIShare homepage for a full platform overview. Registration is free, and you can stop at any time.

Every outbound link in this guide carries a UTM parameter (utm_source=apishare_devto&utm_medium=referral&utm_campaign=free_api_article) for clean campaign attribution.


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free API Cost and Quota Control in Practice: 429 Backoff, RPM Budgets, and Multi-Model Fallback2026 Free OneAPI Unified Gateway: Connect 100+ LLM APIs at Zero Cost in One GuideRun a 550B-Parameter Model for Free: 2026 Nemotron 3 Ultra Complete Guide (OpenRouter Free Tier Tested)2026 Free Embedding Vector Model API Panorama: BGE-M3 / Voyage / Nomic / Google / Azure and 6 Options Tested (September Update)Free Function Calling / Tool Use API Tutorial: DeepSeek / Gemini / Qwen — Zero-Cost Agent Tooling (2026-09-16 Verified)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.