← Back to articles
Unified API Calling

Groq Channel Now Live on APIShare: 13 Free Models (6 Chat + 7 Special-Purpose)

TL;DR

Starting August 22, 2026, APIShare has officially onboarded the Groq channel. Groq's in-house LPU inference hardware is famous for raw speed, and this launch brings 13 free models: 6 chat models for the main rankings + 7 special-purpose models — all through APIShare's unified OpenAI-compatible gateway with a single API key.

Item Details
Channel Groq (LPU inference cloud, ultra-low latency)
Protocol OpenAI-compatible /v1/chat/completions
Total models 13 = 6 chat + 7 special-purpose
Cost Free (subject to per-model rate limits)
Verified status 11 live & tested; 2 TTS models coming soon

1. Chat Models · 6

All of the following were verified working via the chat-completions endpoint:

# Model ID Context Max Output Capabilities
1 openai/gpt-oss-120b 131,072 65,536 Reasoning + tools + JSON mode + structured outputs
2 openai/gpt-oss-20b 131,072 65,536 Reasoning + tools
3 qwen/qwen3.6-27b 131,072 16,384 Multimodal input (text+image) + tools
4 groq/compound 131,072 8,192 Agentic system with built-in tool orchestration
5 groq/compound-mini 131,072 8,192 Lightweight agentic system
6 allam-2-7b 4,096 4,096 Arabic-optimized, blazing fast

Highlights:

  • openai/gpt-oss-120b — OpenAI's open-weight 120B flagship: 131K context, 65K max output, reasoning and structured outputs. The strongest all-around model in this batch.
  • openai/gpt-oss-20b — The lighter sibling; measured near 1,000 tok/s generation. Great for high-frequency calls.
  • qwen/qwen3.6-27b — The only chat model in this batch supporting image input. Go-to choice for visual Q&A and screenshot understanding.
  • groq/compound / compound-mini — Groq's official agentic systems that auto-route underlying models and tools; great for search-augmented tasks. Note stricter rate limits under sustained load.
  • allam-2-7b — Built by Saudi Arabia's SDAIA, deeply optimized for Arabic; trades a small 4K context for extreme speed (~1,260 tok/s).

Note: No Llama chat series (e.g., llama-3.1-8b-instant) is included on this channel — treat the table above as authoritative.

2. Special-Purpose Models · 7

Per our editorial policy, classifiers and audio models are listed separately and excluded from the chat rankings:

🛡️ Content Safety Classification (3)

Model ID Purpose
openai/gpt-oss-safeguard-20b Content safety review (safety-specialized gpt-oss-20b)
meta-llama/llama-prompt-guard-2-86m Prompt injection detection with risk scores
meta-llama/llama-prompt-guard-2-22m Lighter version, lower latency

Ideal as a front-line filter for RAG pipelines and user input screening. Called via chat-completions.

🎙️ Speech-to-Text (2, live)

Model ID Notes
whisper-large-v3 High-accuracy multilingual transcription
whisper-large-v3-turbo Speed-optimized variant

Served via the /audio/transcriptions endpoint.

🔊 Text-to-Speech (2, coming soon)

Model ID Language
canopylabs/orpheus-v1-english English
canopylabs/orpheus-arabic-saudi Arabic (Saudi)

⚠️ These two require Groq org admin to accept model terms in the console before activation. Calls currently return model_terms_required. Do not rely on them in production until activated — we'll update this article once they go live.

3. Quick Start

Register on APIShare and create an API key to call everything through one gateway:

curl https://apishare.cc/v1/chat/completions \
  -H "Authorization: Bearer $APISHARE_API_KEY" \
  -H "Content-Type: application/json" -d '{
    "model": "openai/gpt-oss-120b",
    "messages": [{"role": "user", "content": "Explain LPU vs GPU in three sentences"}],
    "max_tokens": 512
  }'
from openai import OpenAI

client = OpenAI(
    base_url="https://apishare.cc/v1",
    api_key="YOUR_APISHARE_KEY",
)

resp = client.chat.completions.create(
    model="qwen/qwen3.6-27b",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What's in this image?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/demo.png"}}
        ]
    }],
)
print(resp.choices[0].message.content)

Speech-to-text example:

with open("audio.mp3", "rb") as f:
    resp = client.audio.transcriptions.create(
        model="whisper-large-v3-turbo",
        file=f,
    )
print(resp.text)

4. Measured Performance Reference

Single-round sampling (max_tokens=100, temperature=0, measured 2026-08-22):

Model Generation speed End-to-end latency
allam-2-7b ~1264 tok/s 0.19s
openai/gpt-oss-20b ~988 tok/s 0.52s
qwen/qwen3.6-27b ~508 tok/s 0.41s
openai/g件-oss-120b ~486 tok/s 0.67s
groq/compound ~476 tok/s 1.69s*
groq/compound-mini ~455 tok/s —

* compound is a multi-step agentic system; longer queueing is normal.

5. Usage Notes

  1. Reasoning models need headroom: gpt-oss models are reasoning models — set max_tokens ≥256 or answers may be cut off by hidden reasoning tokens.
  2. compound rate limits are strict: the agentic system auto-routes underlying models; sustained load can trigger org-level throttling (~30s recovery). Add retry/backoff in production.
  3. Orpheus TTS not yet active: see section 2 — calls error out until terms are accepted.
  4. No Llama chat series: only the two prompt-guard classifiers carry the meta-llama prefix; no llama-3.x chat models.

Wrap-up

Groq fills the "blazing-fast inference" gap in APIShare's free lineup: run gpt-oss-120b for heavy lifting, switch to gpt-oss-20b for throughput, hand vision tasks to qwen3.6-27b, and front-load safety checks with prompt-guard — all with a single key.

View the Full Groq Lineup on Apishare.cc

This article covers the Groq channel only. To compare free quotas and latency across Groq, OpenRouter and NVIDIA NIM, browse the APIShare free API directory.

To switch models across inference clouds with one key: sign up for Apishare.cc and claim free credits.

Compare Free Inference Clouds Beyond Groq

To compare Groq against OpenRouter, NVIDIA NIM and others, filter models and quotas on https://apishare.cc/free-api?utm_source=article&utm_medium=referral&utm_campaign=b1a_20260929 .

Unified auth and routing: https://apishare.cc/register?utm_source=article&utm_medium=referral&utm_campaign=b1a_20260929 .

Want a side-by-side comparison? Browse the APIShare free API catalog, filter by category, and register to claim a key.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App — Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling — Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool — Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models — Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models — OpenAI Compatible + Smart Routing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.