← Back to articles
Unified API Calling

Cherry Studio Complete Guide: 300+ Models in One Desktop App — Local KB + MCP, Zero-Cost Unified Calling

Cherry Studio Complete Guide: 300+ Models in One Desktop App — Local KB + MCP, Zero-Cost Unified Calling

TL;DR: Cherry Studio is an open-source desktop AI assistant (Electron, 300+ providers) that collapses OpenAI, Claude, Gemini, DeepSeek, Groq, OpenRouter into one UI, with local knowledge base, MCP tools and multi-model side-by-side comparison — bringing unified calling to the desktop.

Why Cherry Studio?

  • 300+ providers: Templates for OpenAI, Anthropic, Gemini, DeepSeek, Groq, Together, OpenRouter, Ollama — paste key and go.
  • Local RAG: Import PDF/Word/Markdown/Excel, vector search with citations, offline.
  • MCP: Native Model Context Protocol — filesystem, web search, code execution pluggable.
  • Multi-model: Send one prompt to 3-5 models in parallel, compare quality/latency.
  • Local-first: SQLite + encrypted keys, no third-party relay.

Setup

Download from https://cherry-ai.com or GitHub CherryHQ/cherry-studio (Windows/macOS/Linux). Settings → Model Services → Add → pick provider → paste Key + Base URL → Check → select models. Top switcher then swaps models instantly.

OpenRouter unified entry
Type: OpenAI Compatible
Base URL: https://openrouter.ai/api/v1
Key: sk-or-v1-xxx
Models: deepseek/deepseek-r1:free, qwen/qwen3-235b:free, moonshotai/kimi-k2:free

Unified Calling

  • Side-by-side: Enable multi-model, pick 3 models, send once, compare.
  • Knowledge Base: Create KB → drag docs → @KB in chat → retrieval-augmented answer with citations.
  • MCP: Settings → MCP → add servers (filesystem, brave-search, python) → model auto-calls tools.

vs OneAPI / LiteLLM

Solution Form Layer Best For
Cherry Studio Desktop Client aggregation Personal multi-model + local RAG
OneAPI Gateway Single binary Server aggregation Team self-host + monetize
LiteLLM Proxy Python proxy Server aggregation Python teams, 100+ models
Portkey Enterprise gateway Server + governance Enterprise guardrails

Best Practices

  • Pin OpenRouter :free models on top for daily free use.
  • Alias models to short names (r1-free, k2).
  • Backup config.json for migration.
  • Connect Ollama at localhost:11434 for offline unified UI.

Summary

  • Free multi-model: Cherry + OpenRouter :free, compare 10+ models after one install.
  • Local KB: Drag-and-drop RAG, no vector DB setup.
  • Gateway link: Point Cherry's Base URL to OneAPI http://localhost:3000/v1 to reuse mymodel routing.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Lobe Chat Complete Guide: Pluginized Web Unified Calling — Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool — Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models — Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models — OpenAI Compatible + Smart RoutingNextChat Complete Guide: Lightweight Web Unified Calling — One-Click Model Switch + Prompt Marketplace

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.