← Back to articles
Unified API Calling

Streaming Response Unified Handling

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Background

OpenAI uses text/event-stream SSE pushing data: {...}\n\n chunks; Anthropic uses event: content_block_delta events; Gemini uses chunked JSON arrays. If clients implement each separately, every new provider forces another rewrite of the streaming parser. The unified gateway translates every upstream streaming protocol into OpenAI SSE so one client parser eats all models.

Core Problems

  • Time-to-first-token (TTFT): users are highly sensitive to waits above 1s. Treat TTFT as a routing metric and prefer fast-start providers (e.g. Groq).
  • Mid-stream failure: if the upstream returns 5xx halfway through, the client has already started rendering — you cannot just retry. The gateway must heartbeat between chunks and, on failure, emit a data: {"error": ...} chunk to close the stream gracefully.
  • Backpressure: upstream pushes faster than the client reads; a buffer overflow drops chunks. The gateway uses asyncio.Queue(maxsize=64) as a bounded queue and pauses upstream reads when full.
  • Cancel propagation: when the client disconnects, the gateway must immediately cancel the upstream request — otherwise the upstream keeps running and burning tokens.

Code Example

import asyncio, json
from fastapi import FastAPI
from fastapi.responses import StreamingResponse

app = FastAPI()

async def sse_translate(upstream_stream, model):
    'Translate any upstream stream into OpenAI SSE.'
    queue = asyncio.Queue(maxsize=64)
    async def producer():
        try:
            async for chunk in upstream_stream:
                # normalize each chunk to OpenAI shape
                data = {
                    "id": "chatcmpl-x", "object": "chat.completion.chunk",
                    "model": model,
                    "choices": [{"index": 0, "delta": {"content": chunk.text},
                                 "finish_reason": None}]}
                await queue.put(f"data: {json.dumps(data)}\n\n")
        except Exception as e:
            await queue.put(f"data: {json.dumps({'error': str(e)})}\n\n")
        finally:
            await queue.put("data: [DONE]\n\n")

    task = asyncio.create_task(producer())
    try:
        while True:
            item = await queue.get()
            if item == "data: [DONE]\n\n":
                break
            yield item
    finally:
        task.cancel()  # propagate cancellation upstream

@app.get("/v1/chat")
async def chat():
    return StreamingResponse(sse_translate(..., "llama-3.3-70b"),
                              media_type="text/event-stream")

Edge Buffering and CDN Coordination

When clients are on cross-border networks, direct gateway connections suffer high latency and packet loss. Routing SSE responses through a CDN (e.g. Cloudflare) edge node buffer dramatically lowers TTFT. But CDNs impose SSE timeouts (usually 100s); long streams must emit heartbeats to keep the connection alive. The gateway can emit : keep-alive comment chunks every 15 seconds — the CDN will not drop the connection, and clients ignore comment lines during parsing. This CDN coordination lets free models serve global users too.

Best Practices

  • Heartbeat: every 15s emit an SSE comment : ping\n\n to keep CDN/reverse proxies from closing the connection.
  • Resume support: implement Last-Event-ID so the client can resume from a break; the gateway caches the last N chunks for this.
  • First-chunk flush: flush the first chunk immediately — do not wait to accumulate 4KB, otherwise TTFT balloons at the proxy layer.
  • Cancel propagation: when the client disconnects, the upstream request must be cancelled to avoid burning tokens on dead streams.

Free models can also deliver a smooth typewriter experience — as long as the gateway aligns every streaming protocol to SSE.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App — Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling — Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool — Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models — Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models — OpenAI Compatible + Smart Routing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.