← Back to articles
Detailed Usage

DeepSeek API Integration Steps

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Introduction

DeepSeek is a top-tier open-source LLM team. Its deepseek-chat (V3) and deepseek-reasoner (R1) excel at reasoning and coding tasks, with token prices far below GPT-4o. This article shows how to integrate them, covering standard chat, the reasoning model, streaming, and function calling.

架构图

flowchart TD A[Register at platform.deepseek.com] --> B[Verify phone] B --> C[Top up ¥10 or use bonus] C --> D[Create API Key] D --> E[Use OpenAI SDK] E --> F{Choose model} F -->|deepseek-chat| G[V3: general chat] F -->|deepseek-reasoner| H[R1: math/code reasoning]

Get an API Key

  1. Visit https://platform.deepseek.com, register, and verify your phone number.
  2. Go to API Keys → Create API Key and copy the sk-... value.
  3. Top up at least ¥10 in Billing (or use the welcome credit).

Install and Configure

DeepSeek is fully OpenAI-compatible, so reuse the openai SDK:

pip install openai
export DEEPSEEK_API_KEY="sk-..."

Standard Chat (deepseek-chat)

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEEPSEEK_API_KEY"],
    base_url="https://api.deepseek.com/v1",
)

resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Explain an inverted index in one sentence."},
    ],
    temperature=0.3,
    max_tokens=256,
)
print(resp.choices[0].message.content)

Reasoning Model (deepseek-reasoner)

R1 returns a chain-of-thought — great for math, logic, and complex code:

resp = client.chat.completions.create(
    model="deepseek-reasoner",
    messages=[
        {"role": "user", "content": "A pool has two inlet pipes A (3h to fill) and B (6h to fill) and an outlet C (4h to drain). How long to fill with all three open?"},
    ],
)
print("Reasoning:")
print(resp.choices[0].message.reasoning_content)
print("\nFinal answer:")
print(resp.choices[0].message.content)

Note: reasoning_content is R1-specific; V3 does not return it.

Streaming

stream = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user", "content": "Write a haiku about autumn."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)
print()

Function Calling

tools = [{
    "type": "function",
    "function": {
        "name": "get_stock_price",
        "description": "Get the real-time price of a stock",
        "parameters": {
            "type": "object",
            "properties": {"symbol": {"type": "string"}},
            "required": ["symbol"],
        },
    },
}]
resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user", "content": "What's the price of AAPL?"}],
    tools=tools,
)
print(resp.choices[0].message.tool_calls[0].function)

Pricing (reference)

  • deepseek-chat: ¥0.5 / 1M input (¥0.1 cache hit), ¥8 / 1M output
  • deepseek-reasoner: ¥4 / 1M input, ¥16 / 1M output (non-cache)

About an order of magnitude cheaper than GPT-4o. Prefix caching saves another 80% for repeated system prompts.

Context Caching

DeepSeek supports context caching for repeated system prompts — cache hits drop input price to ¥0.1/1M (from ¥0.5). This is great for multi-turn agents and RAG lookups. Caching is automatic, but the prompt prefix must be byte-identical. Put fixed system instructions and tool definitions at the start of the messages array to maximize hit rate.

Troubleshooting

  • 402 Insufficient Balance: Out of credit. Top up in Billing.
  • R1 output too long: Cap with max_tokens or truncate reasoning_content.
  • Self-host: Download the open weights and serve with vLLM or SGLang. R1 distill fits on 2x A100.
  • Slow responses from abroad: Latency may be high. Use a proxy or the DeepSeek mirror on OpenRouter.

DeepSeek is one of the most cost-effective choices for Chinese-language workloads.

Best Practices

  • Use deepseek-reasoner for reasoning tasks: R1 for math/code review/logic; V3 for daily chat saves tokens.
  • Pricing is per-token: input 0.27 yuan/M, output 1.1 yuan/M (CNY) — 100x cheaper than GPT-4o.
  • Always enable streaming: reasoner takes long to think; without streaming users think it's stuck.
  • Never hardcode API Key in frontend: DeepSeek bills per key; frontend hardcode is giving money away.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free AI Content Moderation API Guide 2026: Llama Guard 3 vs Perspective vs OpenAIFree OCR and Document Parsing API in PracticeIntegrating Free APIs into Your Local IDEConnecting Free Models to OpenCode in PracticeApplying for an OpenRouter API Key and Understanding Pricing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.