⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Introduction
Groq is a hardware company focused on inference acceleration. Its LPU (Language Processing Unit) runs Llama-family models at 500+ tokens/s — one to two orders of magnitude faster than typical GPUs. Better yet, it offers a generous free tier and is fully OpenAI-compatible. This article gets you connected in three steps and demos streaming and tool use.
架构图
Step 1: Get an API Key
- Visit https://console.groq.com and sign in with Google or GitHub.
- Go to API Keys → Create API Key and copy the string starting with
gsk_.... - Export it:
export GROQ_API_KEY="gsk_..."
Step 2: Install the SDK
Groq maintains an official groq Python SDK. You can also use the openai SDK, but the former has better type hints:
pip install groq
Step 3: Make a Call
import os
from groq import Groq
client = Groq(api_key=os.environ["GROQ_API_KEY"])
resp = client.chat.completions.create(
model="llama-3.3-70b-versatile",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain an inverted index in one sentence."},
],
temperature=0.3,
max_tokens=256,
)
print(resp.choices[0].message.content)
Streaming
The best way to feel LPU speed is streaming — first-token latency is often under 200ms:
stream = client.chat.completions.create(
model="llama-3.3-70b-versatile",
messages=[{"role": "user", "content": "Write a haiku about autumn."}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
print()
Tool Use
Groq supports OpenAI-style function calling, ideal for agent workflows:
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
resp = client.chat.completions.create(
model="llama-3.3-70b-versatile",
messages=[{"role": "user", "content": "What's the weather in Tokyo today?"}],
tools=tools,
)
call = resp.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)
# get_weather {"city": "Tokyo"}
Free-tier Limits (as of 2025)
- Requests per minute: 30 RPM (70B), 1440 RPM (8B)
- Daily token cap: ~1M tokens/day, varying by model
- Context window: 128K for Llama 3.3 70B
Real-world Speed Benchmarks
On Llama-3.3-70B, Groq's LPU sustains 250-500 tokens/s per request, whereas the same model on typical GPU inference services only reaches 30-80 tokens/s. For latency-sensitive use cases — real-time chat, code completion, multi-turn agents — this gap translates directly into better UX. Note that for long inputs (>8K tokens) the LPU's advantage narrows in the prefill phase, where GPU memory bandwidth wins.
Troubleshooting
429 rate_limit_exceeded: Free-tier cap hit. Wait 60 seconds or switch to the smallerllama-3.1-8b-instantmodel.- JSON mode: Add
response_format={"type":"json_object"}and prompt the model to output JSON. - List available models: Visit https://api.groq.com/openai/v1/models.
- Truncated responses:
max_tokensdefaults low — set it to 1024+.
Groq is ideal for latency-sensitive use cases like real-time chat and code completion.
Best Practices
- Use batch requests: Groq supports batch (one request, multiple turns) — 3-5x throughput.
- Stay below 30 RPM: implement a local token bucket at 25 RPM to leave buffer.
- Prefer streaming: 500+ tps streaming UX beats waiting for the full response.
- Mixtral 8x7B is fastest on Groq: MoE models benefit more from LPU acceleration than dense models.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key