GLM-4.7 Flash Free API Tutorial: 128K Ultra-Fast Inference Forever Free
Updated: 2026-09-02 · Official: https://www.zhipuai.cn · OpenRouter: https://openrouter.ai/models/z-ai/glm-4.7-flash:free · Docs: https://docs.z.ai · Verified: 2026-09-02 Dual-route + ratelimit headers
GLM-4.7 Flash is Zhipu's 2025 flagship lightweight MoE — 30B total 3B active, 128K context, ultra-fast Flash inference. Free via z-ai/glm-4.7-flash:free (20 RPM) forever $0, OpenAI compatible.
Why GLM-4.7 Flash (5D Score 22/25)
| Dimension | GLM-4.7 Flash | Score |
|---|---|---|
| Free | 20 RPM forever $0, $0.1/1M paid | 5 |
| Stability | Production SLA | 4 |
| Latency | p50 <300ms | 5 |
| Rate Limit | 20 RPM / 50 req/day | 4 |
| Intelligence | MMLU 86%+ | 4 |
| Model | Free ($/1M) | Stability | Latency p50 | Limit RPM | Intel | Total |
|---|---|---|---|---|---|---|
| GLM-4.7 Flash | $0/$0 forever | 4 | 0.30s | 20 | 86%/88% | 22 |
| DeepSeek-V3 | $0.55/$2.19 | 5 | 1.2s | 100 | 84% | 20 |
| Kimi K2 | $0.6/$2.2 | 4 | 0.5s | 60/20 | 88% | 21 |
Quick Start
curl -X POST https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-d '{"model":"z-ai/glm-4.7-flash:free","messages":[{"role":"user","content":"hi"}]}'
curl -i -X POST https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-d '{"model":"z-ai/glm-4.7-flash:free","messages":[{"role":"user","content":"hi"}]}' | head -n 30
from openai import OpenAI
client = OpenAI(api_key="OR_KEY", base_url="https://openrouter.ai/api/v1")
print(client.chat.completions.create(model="z-ai/glm-4.7-flash:free", messages=[{"role":"user","content":"hi"}]).choices[0].message.content)
import OpenAI from "openai";
const c = new OpenAI({apiKey: process.env.OPENROUTER_API_KEY, baseURL:"https://openrouter.ai/api/v1"});
console.log((await c.chat.completions.create({model:"z-ai/glm-4.7-flash:free", messages:[{role:"user",content:"hi"}]})).choices[0].message.content);
4) Ratelimit Header Screenshots (Required)

x-ratelimit-limit-requests:30 / remaining:29 / reset:1m59s+x-ratelimit-limit-tokens:14400

x-ratelimit-limit:20 / remaining:49 / reset:5m37s+x-or-ratelimit-limit:20+ HTTP 200
Verified 200
GET /v1/models200,POST /chat/completions200, context 131072
Sources
Updated: 2026-09-02 · Price/limits refresh within 24h
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
-
Full model catalog: APIShare free API directory
-
Sign up for a free trial key: Register and claim your API key