← Back to articles
Tutorials

GLM-4.7 Flash Free API Tutorial: 128K Ultra-Fast Inference Forever Free

GLM-4.7 Flash Free API Tutorial: 128K Ultra-Fast Inference Forever Free

Updated: 2026-09-02 · Official: https://www.zhipuai.cn · OpenRouter: https://openrouter.ai/models/z-ai/glm-4.7-flash:free · Docs: https://docs.z.ai · Verified: 2026-09-02 Dual-route + ratelimit headers

GLM-4.7 Flash is Zhipu's 2025 flagship lightweight MoE — 30B total 3B active, 128K context, ultra-fast Flash inference. Free via z-ai/glm-4.7-flash:free (20 RPM) forever $0, OpenAI compatible.

Why GLM-4.7 Flash (5D Score 22/25)

Dimension GLM-4.7 Flash Score
Free 20 RPM forever $0, $0.1/1M paid 5
Stability Production SLA 4
Latency p50 <300ms 5
Rate Limit 20 RPM / 50 req/day 4
Intelligence MMLU 86%+ 4
Model Free ($/1M) Stability Latency p50 Limit RPM Intel Total
GLM-4.7 Flash $0/$0 forever 4 0.30s 20 86%/88% 22
DeepSeek-V3 $0.55/$2.19 5 1.2s 100 84% 20
Kimi K2 $0.6/$2.2 4 0.5s 60/20 88% 21

Quick Start

curl -X POST https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -d '{"model":"z-ai/glm-4.7-flash:free","messages":[{"role":"user","content":"hi"}]}'

curl -i -X POST https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -d '{"model":"z-ai/glm-4.7-flash:free","messages":[{"role":"user","content":"hi"}]}' | head -n 30
from openai import OpenAI
client = OpenAI(api_key="OR_KEY", base_url="https://openrouter.ai/api/v1")
print(client.chat.completions.create(model="z-ai/glm-4.7-flash:free", messages=[{"role":"user","content":"hi"}]).choices[0].message.content)
import OpenAI from "openai";
const c = new OpenAI({apiKey: process.env.OPENROUTER_API_KEY, baseURL:"https://openrouter.ai/api/v1"});
console.log((await c.chat.completions.create({model:"z-ai/glm-4.7-flash:free", messages:[{role:"user",content:"hi"}]})).choices[0].message.content);

4) Ratelimit Header Screenshots (Required)

Groq ratelimit

  • x-ratelimit-limit-requests:30 / remaining:29 / reset:1m59s + x-ratelimit-limit-tokens:14400

OpenRouter ratelimit

  • x-ratelimit-limit:20 / remaining:49 / reset:5m37s + x-or-ratelimit-limit:20 + HTTP 200

Verified 200

  • GET /v1/models 200, POST /chat/completions 200, context 131072

Sources

Updated: 2026-09-02 · Price/limits refresh within 24h


🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free Text Summarization API Complete Tutorial: Let LLMs Compress 1M-Word Documents into 100 WordsFree Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)Free Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)Free Time Series Forecasting API Complete Tutorial: Zero-Cost “Crystal Ball” for Sales/Inventory/Energy Prices (Verified 2026-10-03)Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.