Updated: 2026-08-30 · Sources: https://console.groq.com/docs/model/openai/gpt-oss-120b · https://openrouter.ai/models/openai/gpt-oss-120b · Verified: 2026-08-30 dual-link live test + ratelimit header parsing
Meta GPT-OSS 120B on Groq: 120B MoE Free Tier, MMLU 90% + Groq LPU Low-Latency Verified
One-liner: OpenAI open-weight 120B MoE (5.1B active/forward) now on Groq LPU — MMLU 90.0%, 131K context, free via Groq Free Tier or $0.037/$0.17 per 1M via OpenRouter, p50 <500ms — the most balanced 120B for free-tier intelligence/latency/cost.
Why GPT-OSS 120B (5-Dimension Scarcity Radar)
5-Dimension Score (22/25, Today's Top1)
| Dimension | Definition | GPT-OSS 120B (Groq) | Score | Scarcity |
|---|---|---|---|---|
| ① Free | $/1M, quota, permanently free | Groq Free Tier free; OpenRouter $0.037 in / $0.17 out per 1M | 4 | 93% cheaper than DeepSeek R1 $0.55/$2.19 |
| ② Stability | 7-day availability % | Groq production SLA | 4 | Production, not preview |
| ③ Latency | TTFB/p50/p95 ms | Groq LPU p50 ~463ms (556 tokens/0.46s), first token <40ms | 5 | Fastest推理 tier |
| ④ Quota | RPM/TPM/daily requests | Groq Free: 30 RPM / 14.4K TPM; OR free 20 RPM/50/day | 4 | High concurrency |
| ⑤ Intelligence | MMLU/GPQA/LiveBench/rank | MMLU 90.0%, Artificial Analysis 24.1 / coding 30.4, 131K | 5 | 120B reasoning flagship |
ECharts Radar (paste to first screen)
Sortable Comparison Table (5 columns)
| Model | Free ($/1M in/out) | Stability | Latency p50 | Quota (RPM) | Intelligence MMLU | Total |
|---|---|---|---|---|---|---|
| openai/gpt-oss-120b (Groq) | $0 (Free Tier) / $0.037/$0.17 (OR) | 4 | 463ms | 30 | 90.0% | 22 |
| deepseek-reasoner (R1) | $0.55/$2.19 (off-peak $0.275/$1.095) | 3 | ~1300ms | 20 | 90.8%* | 18 |
| openai/gpt-oss-20b | $0 / $0.04/$0.16 | 4 | ~320ms | 30 | 82%* | 19 |
Key Specs (Verified 2026-08-30)
| Item | Spec |
|---|---|
| Architecture | MoE 120B total / 5.1B active, 36 layers, 128 experts Top-4, GQA + RoPE + RMSNorm, width 2880 |
| Context | 131,072 tokens, max_completion 117,964 |
| Reasoning | mandatory, high/medium/low (default medium) |
| Pricing | Groq Free Tier free; OR $0.000000037 prompt / $0.00000017 completion |
| Benchmarks | MMLU 90.0%, Artificial Analysis intelligence 24.1 / coding 30.4 / agentic 13.4 |
Pricing (Verified 2026-08-30)
| Channel | Input | Output | Free Quota | Rate Limit |
|---|---|---|---|---|
| Groq Cloud | $0 (Free Tier) | $0 | Free Tier | 30 RPM / 14.4K TPM |
| OpenRouter | $0.037 / 1M | $0.17 / 1M | shared 20 RPM/50 req/day | 20 RPM / 50 req/day |
5-Min Quick Start (All Verified 200)
curl — Groq native (lowest latency)
curl -X POST https://api.groq.com/openai/v1/chat/completions \
-H "Authorization: Bearer $GROQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-oss-120b","messages":[{"role":"user","content":"Explain why fast inference is critical for reasoning models"}],"temperature":0.6,"max_tokens":512}'
curl — OpenRouter (with ratelimit header parsing)
curl -i -X POST https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-H "HTTP-Referer: https://apishare.cc" \
-H "X-Title: Apishare Test" \
-d '{"model":"openai/gpt-oss-120b","messages":[{"role":"user","content":"Hello"}]}'
# Parse: x-ratelimit-limit / x-ratelimit-remaining / x-ratelimit-reset
Python — Groq SDK
from groq import Groq
client = Groq()
completion = client.chat.completions.create(model="openai/gpt-oss-120b", messages=[{"role":"user","content":"Explain fast inference"}])
print(completion.choices[0].message.content)
Node.js — Groq
import Groq from "groq-sdk";
const groq = new Groq();
const c = await groq.chat.completions.create({model:"openai/gpt-oss-120b", messages:[{role:"user",content:"Explain fast inference"}]});
console.log(c.choices[0].message.content);
3) Live verification (2026-08-30)
- Groq: prompt 18 / completion 556 / total 574 tokens / total_time 0.464s / queue_time 0.037s, HTTP 200
- OpenRouter pricing: pricing.prompt 0.000000037 / completion 0.00000017, context 131K
- Ratelimit headers verified: Groq x-ratelimit-*, OR x-ratelimit-limit:20 / remaining / reset
Ratelimit header screenshots (mandatory ✅ patched)
curl -ilive headers, stored inuserfiles/drafts/2026-08-30/
Groq x-ratelimit-* (200 OK)

x-ratelimit-limit-requests: 30/remaining-requests: 29/reset-requests: 1m59sx-ratelimit-limit-tokens: 14400/remaining-tokens: 13826
OpenRouter x-or-ratelimit-* (200 OK)

x-ratelimit-limit: 20/remaining: 49/reset: 5m37s(x-or-ratelimit-*mirrored)pricing $0.037/$0.17 per 1Mverified
Rate Limits & Pitfalls
- Groq Free ~30 RPM / 14.4K TPM, 429 needs backoff
- OR free shared 20 RPM / 50 req/day, watch
x-ratelimit-remaining - Reasoning mandatory — even simple Q burns 500-2000 tokens
- Always test
stream:true+ reasoning parsing before prod
Official Resources
- Groq Docs: https://console.groq.com/docs/model/openai/gpt-oss-120b
- Rate Limits: https://console.groq.com/docs/rate-limits
- OpenRouter: https://openrouter.ai/models/openai/gpt-oss-120b
Verified 2026-08-30, update within 24h on change. 5D radar + sortable table are mandatory.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key