TL;DR
Starting August 22, 2026, APIShare has officially onboarded the Groq channel. Groq's in-house LPU inference hardware is famous for raw speed, and this launch brings 13 free models: 6 chat models for the main rankings + 7 special-purpose models — all through APIShare's unified OpenAI-compatible gateway with a single API key.
| Item | Details |
|---|---|
| Channel | Groq (LPU inference cloud, ultra-low latency) |
| Protocol | OpenAI-compatible /v1/chat/completions |
| Total models | 13 = 6 chat + 7 special-purpose |
| Cost | Free (subject to per-model rate limits) |
| Verified status | 11 live & tested; 2 TTS models coming soon |
1. Chat Models · 6
All of the following were verified working via the chat-completions endpoint:
| # | Model ID | Context | Max Output | Capabilities |
|---|---|---|---|---|
| 1 | openai/gpt-oss-120b |
131,072 | 65,536 | Reasoning + tools + JSON mode + structured outputs |
| 2 | openai/gpt-oss-20b |
131,072 | 65,536 | Reasoning + tools |
| 3 | qwen/qwen3.6-27b |
131,072 | 16,384 | Multimodal input (text+image) + tools |
| 4 | groq/compound |
131,072 | 8,192 | Agentic system with built-in tool orchestration |
| 5 | groq/compound-mini |
131,072 | 8,192 | Lightweight agentic system |
| 6 | allam-2-7b |
4,096 | 4,096 | Arabic-optimized, blazing fast |
Highlights:
- openai/gpt-oss-120b — OpenAI's open-weight 120B flagship: 131K context, 65K max output, reasoning and structured outputs. The strongest all-around model in this batch.
- openai/gpt-oss-20b — The lighter sibling; measured near 1,000 tok/s generation. Great for high-frequency calls.
- qwen/qwen3.6-27b — The only chat model in this batch supporting image input. Go-to choice for visual Q&A and screenshot understanding.
- groq/compound / compound-mini — Groq's official agentic systems that auto-route underlying models and tools; great for search-augmented tasks. Note stricter rate limits under sustained load.
- allam-2-7b — Built by Saudi Arabia's SDAIA, deeply optimized for Arabic; trades a small 4K context for extreme speed (~1,260 tok/s).
Note: No Llama chat series (e.g., llama-3.1-8b-instant) is included on this channel — treat the table above as authoritative.
2. Special-Purpose Models · 7
Per our editorial policy, classifiers and audio models are listed separately and excluded from the chat rankings:
🛡️ Content Safety Classification (3)
| Model ID | Purpose |
|---|---|
openai/gpt-oss-safeguard-20b |
Content safety review (safety-specialized gpt-oss-20b) |
meta-llama/llama-prompt-guard-2-86m |
Prompt injection detection with risk scores |
meta-llama/llama-prompt-guard-2-22m |
Lighter version, lower latency |
Ideal as a front-line filter for RAG pipelines and user input screening. Called via chat-completions.
🎙️ Speech-to-Text (2, live)
| Model ID | Notes |
|---|---|
whisper-large-v3 |
High-accuracy multilingual transcription |
whisper-large-v3-turbo |
Speed-optimized variant |
Served via the /audio/transcriptions endpoint.
🔊 Text-to-Speech (2, coming soon)
| Model ID | Language |
|---|---|
canopylabs/orpheus-v1-english |
English |
canopylabs/orpheus-arabic-saudi |
Arabic (Saudi) |
⚠️ These two require Groq org admin to accept model terms in the console before activation. Calls currently return model_terms_required. Do not rely on them in production until activated — we'll update this article once they go live.
3. Quick Start
Register on APIShare and create an API key to call everything through one gateway:
curl https://apishare.cc/v1/chat/completions \
-H "Authorization: Bearer $APISHARE_API_KEY" \
-H "Content-Type: application/json" -d '{
"model": "openai/gpt-oss-120b",
"messages": [{"role": "user", "content": "Explain LPU vs GPU in three sentences"}],
"max_tokens": 512
}'
from openai import OpenAI
client = OpenAI(
base_url="https://apishare.cc/v1",
api_key="YOUR_APISHARE_KEY",
)
resp = client.chat.completions.create(
model="qwen/qwen3.6-27b",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What's in this image?"},
{"type": "image_url", "image_url": {"url": "https://example.com/demo.png"}}
]
}],
)
print(resp.choices[0].message.content)
Speech-to-text example:
with open("audio.mp3", "rb") as f:
resp = client.audio.transcriptions.create(
model="whisper-large-v3-turbo",
file=f,
)
print(resp.text)
4. Measured Performance Reference
Single-round sampling (max_tokens=100, temperature=0, measured 2026-08-22):
| Model | Generation speed | End-to-end latency |
|---|---|---|
| allam-2-7b | ~1264 tok/s | 0.19s |
| openai/gpt-oss-20b | ~988 tok/s | 0.52s |
| qwen/qwen3.6-27b | ~508 tok/s | 0.41s |
| openai/g件-oss-120b | ~486 tok/s | 0.67s |
| groq/compound | ~476 tok/s | 1.69s* |
| groq/compound-mini | ~455 tok/s | — |
* compound is a multi-step agentic system; longer queueing is normal.
5. Usage Notes
- Reasoning models need headroom:
gpt-ossmodels are reasoning models — setmax_tokens≥256 or answers may be cut off by hidden reasoning tokens. - compound rate limits are strict: the agentic system auto-routes underlying models; sustained load can trigger org-level throttling (~30s recovery). Add retry/backoff in production.
- Orpheus TTS not yet active: see section 2 — calls error out until terms are accepted.
- No Llama chat series: only the two prompt-guard classifiers carry the meta-llama prefix; no llama-3.x chat models.
Wrap-up
Groq fills the "blazing-fast inference" gap in APIShare's free lineup: run gpt-oss-120b for heavy lifting, switch to gpt-oss-20b for throughput, hand vision tasks to qwen3.6-27b, and front-load safety checks with prompt-guard — all with a single key.
- 📊 Full free model leaderboard: Free LLM API Rankings
- 🔑 Get your API key: Sign up now
View the Full Groq Lineup on Apishare.cc
This article covers the Groq channel only. To compare free quotas and latency across Groq, OpenRouter and NVIDIA NIM, browse the APIShare free API directory.
To switch models across inference clouds with one key: sign up for Apishare.cc and claim free credits.
Compare Free Inference Clouds Beyond Groq
To compare Groq against OpenRouter, NVIDIA NIM and others, filter models and quotas on https://apishare.cc/free-api?utm_source=article&utm_medium=referral&utm_campaign=b1a_20260929 .
Unified auth and routing: https://apishare.cc/register?utm_source=article&utm_medium=referral&utm_campaign=b1a_20260929 .
Want a side-by-side comparison? Browse the APIShare free API catalog, filter by category, and register to claim a key.