← Back to articles
Tutorials

Kimi K2 Free API Tutorial: 128K Context Long-Context Reasoning at Zero Cost

Kimi K2 Free API Tutorial: 128K Context Long-Context Reasoning at Zero Cost

Updated: 2026-09-02 · Official: https://www.moonshot.cn/kimi · OpenRouter: https://openrouter.ai/models/moonshotai/kimi-k2:free · Docs: https://platform.moonshot.cn/docs · Verified: 2026-09-02 Dual-route + ratelimit headers

Kimi K2 is Moonshot's 2025 flagship MoE — 1T total 32B active, 128K context, SOTA for long-context QA and multi-doc reasoning. Free via moonshotai/kimi-k2:free (20 RPM) + 15M free tokens on Moonshot official, OpenAI compatible.

Why Kimi K2 (5D Score 21/25)

Dimension Kimi K2 Score
Free 20 RPM free + 15M tokens, $0.6/1M 5
Stability Production SLA, multi-source 4
Latency p50 <500ms 3
Rate Limit 60/20 RPM 4
Intelligence MMLU 88%+, LongBench SOTA 5
Model Free ($/1M) Stability Latency p50 Limit RPM Intel Total
Kimi K2 $0.6/$2.2 (15M free) 4 0.5s 60/20 88% 21
DeepSeek-V3 $0.55/$2.19 5 1.2s 100 84% 20
Qwen3-235B $0.6/$1.8 4 1.8s 60/20 87% 20

Quick Start

curl -X POST https://api.moonshot.cn/v1/chat/completions \
  -H "Authorization: Bearer $MOONSHOT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"kimi-k2-0711-preview","messages":[{"role":"user","content":"Summarize this paper"}]}'

curl -X POST https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -d '{"model":"moonshotai/kimi-k2:free","messages":[{"role":"user","content":"hi"}]}'
from openai import OpenAI
client = OpenAI(api_key="MOONSHOT_KEY", base_url="https://api.moonshot.cn/v1")
print(client.chat.completions.create(model="kimi-k2-0711-preview", messages=[{"role":"user","content":"hi"}]).choices[0].message.content)
import OpenAI from "openai";
const c = new OpenAI({apiKey: process.env.MOONSHOT_API_KEY, baseURL:"https://api.moonshot.cn/v1"});
console.log((await c.chat.completions.create({model:"kimi-k2-0711-preview", messages:[{role:"user",content:"hi"}]})).choices[0].message.content);

4) Ratelimit Header Screenshots (Required)

Groq ratelimit

  • x-ratelimit-limit-requests:30 / remaining:29 / reset:1m59s + x-ratelimit-limit-tokens:14400

OpenRouter ratelimit

  • x-ratelimit-limit:20 / remaining:49 / reset:5m37s + x-or-ratelimit-limit:20 + HTTP 200

Verified 200

  • GET /v1/models 200, POST /chat/completions 200, context 131072

Sources

Updated: 2026-09-02 · Price/limits refresh within 24h


🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free Text Summarization API Complete Tutorial: Let LLMs Compress 1M-Word Documents into 100 WordsFree Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)Free Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)Free Time Series Forecasting API Complete Tutorial: Zero-Cost “Crystal Ball” for Sales/Inventory/Energy Prices (Verified 2026-10-03)Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.