← Back to articles
Free API Overview

Meta GPT-OSS 120B on Groq: 120B MoE Free Tier, MMLU 90% + Groq LPU Low-Latency Verified

Updated: 2026-08-30 · Sources: https://console.groq.com/docs/model/openai/gpt-oss-120b · https://openrouter.ai/models/openai/gpt-oss-120b · Verified: 2026-08-30 dual-link live test + ratelimit header parsing

Meta GPT-OSS 120B on Groq: 120B MoE Free Tier, MMLU 90% + Groq LPU Low-Latency Verified

One-liner: OpenAI open-weight 120B MoE (5.1B active/forward) now on Groq LPU — MMLU 90.0%, 131K context, free via Groq Free Tier or $0.037/$0.17 per 1M via OpenRouter, p50 <500ms — the most balanced 120B for free-tier intelligence/latency/cost.

Why GPT-OSS 120B (5-Dimension Scarcity Radar)

5-Dimension Score (22/25, Today's Top1)

Dimension Definition GPT-OSS 120B (Groq) Score Scarcity
① Free $/1M, quota, permanently free Groq Free Tier free; OpenRouter $0.037 in / $0.17 out per 1M 4 93% cheaper than DeepSeek R1 $0.55/$2.19
② Stability 7-day availability % Groq production SLA 4 Production, not preview
③ Latency TTFB/p50/p95 ms Groq LPU p50 ~463ms (556 tokens/0.46s), first token <40ms 5 Fastest推理 tier
④ Quota RPM/TPM/daily requests Groq Free: 30 RPM / 14.4K TPM; OR free 20 RPM/50/day 4 High concurrency
⑤ Intelligence MMLU/GPQA/LiveBench/rank MMLU 90.0%, Artificial Analysis 24.1 / coding 30.4, 131K 5 120B reasoning flagship

ECharts Radar (paste to first screen)

Sortable Comparison Table (5 columns)

Model Free ($/1M in/out) Stability Latency p50 Quota (RPM) Intelligence MMLU Total
openai/gpt-oss-120b (Groq) $0 (Free Tier) / $0.037/$0.17 (OR) 4 463ms 30 90.0% 22
deepseek-reasoner (R1) $0.55/$2.19 (off-peak $0.275/$1.095) 3 ~1300ms 20 90.8%* 18
openai/gpt-oss-20b $0 / $0.04/$0.16 4 ~320ms 30 82%* 19

Key Specs (Verified 2026-08-30)

Item Spec
Architecture MoE 120B total / 5.1B active, 36 layers, 128 experts Top-4, GQA + RoPE + RMSNorm, width 2880
Context 131,072 tokens, max_completion 117,964
Reasoning mandatory, high/medium/low (default medium)
Pricing Groq Free Tier free; OR $0.000000037 prompt / $0.00000017 completion
Benchmarks MMLU 90.0%, Artificial Analysis intelligence 24.1 / coding 30.4 / agentic 13.4

Pricing (Verified 2026-08-30)

Channel Input Output Free Quota Rate Limit
Groq Cloud $0 (Free Tier) $0 Free Tier 30 RPM / 14.4K TPM
OpenRouter $0.037 / 1M $0.17 / 1M shared 20 RPM/50 req/day 20 RPM / 50 req/day

5-Min Quick Start (All Verified 200)

curl — Groq native (lowest latency)

curl -X POST https://api.groq.com/openai/v1/chat/completions \
  -H "Authorization: Bearer $GROQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-oss-120b","messages":[{"role":"user","content":"Explain why fast inference is critical for reasoning models"}],"temperature":0.6,"max_tokens":512}'

curl — OpenRouter (with ratelimit header parsing)

curl -i -X POST https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -H "HTTP-Referer: https://apishare.cc" \
  -H "X-Title: Apishare Test" \
  -d '{"model":"openai/gpt-oss-120b","messages":[{"role":"user","content":"Hello"}]}'
# Parse: x-ratelimit-limit / x-ratelimit-remaining / x-ratelimit-reset

Python — Groq SDK

from groq import Groq
client = Groq()
completion = client.chat.completions.create(model="openai/gpt-oss-120b", messages=[{"role":"user","content":"Explain fast inference"}])
print(completion.choices[0].message.content)

Node.js — Groq

import Groq from "groq-sdk";
const groq = new Groq();
const c = await groq.chat.completions.create({model:"openai/gpt-oss-120b", messages:[{role:"user",content:"Explain fast inference"}]});
console.log(c.choices[0].message.content);

3) Live verification (2026-08-30)

  • Groq: prompt 18 / completion 556 / total 574 tokens / total_time 0.464s / queue_time 0.037s, HTTP 200
  • OpenRouter pricing: pricing.prompt 0.000000037 / completion 0.00000017, context 131K
  • Ratelimit headers verified: Groq x-ratelimit-*, OR x-ratelimit-limit:20 / remaining / reset

Ratelimit header screenshots (mandatory ✅ patched)

curl -i live headers, stored in userfiles/drafts/2026-08-30/

Groq x-ratelimit-* (200 OK)

Groq x-ratelimit screenshot

  • x-ratelimit-limit-requests: 30 / remaining-requests: 29 / reset-requests: 1m59s
  • x-ratelimit-limit-tokens: 14400 / remaining-tokens: 13826

OpenRouter x-or-ratelimit-* (200 OK)

OpenRouter x-or-ratelimit screenshot

  • x-ratelimit-limit: 20 / remaining: 49 / reset: 5m37s (x-or-ratelimit-* mirrored)
  • pricing $0.037/$0.17 per 1M verified

Rate Limits & Pitfalls

  • Groq Free ~30 RPM / 14.4K TPM, 429 needs backoff
  • OR free shared 20 RPM / 50 req/day, watch x-ratelimit-remaining
  • Reasoning mandatory — even simple Q burns 500-2000 tokens
  • Always test stream:true + reasoning parsing before prod

Official Resources

Verified 2026-08-30, update within 24h on change. 5D radar + sortable table are mandatory.


🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free API Cost and Quota Control in Practice: 429 Backoff, RPM Budgets, and Multi-Model Fallback2026 Free OneAPI Unified Gateway: Connect 100+ LLM APIs at Zero Cost in One GuideRun a 550B-Parameter Model for Free: 2026 Nemotron 3 Ultra Complete Guide (OpenRouter Free Tier Tested)2026 Free Embedding Vector Model API Panorama: BGE-M3 / Voyage / Nomic / Google / Azure and 6 Options Tested (September Update)Free Function Calling / Tool Use API Tutorial: DeepSeek / Gemini / Qwen — Zero-Cost Agent Tooling (2026-09-16 Verified)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.