← Back to articles
Detailed Usage

Groq Ultra-fast Inference Tutorial

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Introduction

Groq is a hardware company focused on inference acceleration. Its LPU (Language Processing Unit) runs Llama-family models at 500+ tokens/s — one to two orders of magnitude faster than typical GPUs. Better yet, it offers a generous free tier and is fully OpenAI-compatible. This article gets you connected in three steps and demos streaming and tool use.

架构图

flowchart LR A[Sign up at console.groq.com] --> B[Create API Key] B --> C[Install openai SDK] C --> D[Call llama-3.3-70b-versatile] D --> E[500+ tokens/s]

Step 1: Get an API Key

  1. Visit https://console.groq.com and sign in with Google or GitHub.
  2. Go to API Keys → Create API Key and copy the string starting with gsk_....
  3. Export it:
export GROQ_API_KEY="gsk_..."

Step 2: Install the SDK

Groq maintains an official groq Python SDK. You can also use the openai SDK, but the former has better type hints:

pip install groq

Step 3: Make a Call

import os
from groq import Groq

client = Groq(api_key=os.environ["GROQ_API_KEY"])

resp = client.chat.completions.create(
    model="llama-3.3-70b-versatile",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Explain an inverted index in one sentence."},
    ],
    temperature=0.3,
    max_tokens=256,
)
print(resp.choices[0].message.content)

Streaming

The best way to feel LPU speed is streaming — first-token latency is often under 200ms:

stream = client.chat.completions.create(
    model="llama-3.3-70b-versatile",
    messages=[{"role": "user", "content": "Write a haiku about autumn."}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)
print()

Tool Use

Groq supports OpenAI-style function calling, ideal for agent workflows:

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]
resp = client.chat.completions.create(
    model="llama-3.3-70b-versatile",
    messages=[{"role": "user", "content": "What's the weather in Tokyo today?"}],
    tools=tools,
)
call = resp.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)
# get_weather {"city": "Tokyo"}

Free-tier Limits (as of 2025)

  • Requests per minute: 30 RPM (70B), 1440 RPM (8B)
  • Daily token cap: ~1M tokens/day, varying by model
  • Context window: 128K for Llama 3.3 70B

Real-world Speed Benchmarks

On Llama-3.3-70B, Groq's LPU sustains 250-500 tokens/s per request, whereas the same model on typical GPU inference services only reaches 30-80 tokens/s. For latency-sensitive use cases — real-time chat, code completion, multi-turn agents — this gap translates directly into better UX. Note that for long inputs (>8K tokens) the LPU's advantage narrows in the prefill phase, where GPU memory bandwidth wins.

Troubleshooting

  • 429 rate_limit_exceeded: Free-tier cap hit. Wait 60 seconds or switch to the smaller llama-3.1-8b-instant model.
  • JSON mode: Add response_format={"type":"json_object"} and prompt the model to output JSON.
  • List available models: Visit https://api.groq.com/openai/v1/models.
  • Truncated responses: max_tokens defaults low — set it to 1024+.

Groq is ideal for latency-sensitive use cases like real-time chat and code completion.

Best Practices

  • Use batch requests: Groq supports batch (one request, multiple turns) — 3-5x throughput.
  • Stay below 30 RPM: implement a local token bucket at 25 RPM to leave buffer.
  • Prefer streaming: 500+ tps streaming UX beats waiting for the full response.
  • Mixtral 8x7B is fastest on Groq: MoE models benefit more from LPU acceleration than dense models.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free AI Content Moderation API Guide 2026: Llama Guard 3 vs Perspective vs OpenAIFree OCR and Document Parsing API in PracticeIntegrating Free APIs into Your Local IDEConnecting Free Models to OpenCode in PracticeApplying for an OpenRouter API Key and Understanding Pricing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.