← Back to articles
Unified API Calling

Mobile Unified Calling SDK

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Background

Calling models from mobile has three unique problems: network jitter (subway, elevator), battery budget (no sustained background runs), and first-screen latency sensitivity (users tolerate >2s waits poorly). Using openai-python directly does not work — it was written for servers, with no retry backoff, no streaming backpressure, no offline cache. iOS and Android each need a lightweight SDK wrapping the unified gateway.

SDK Design Points

  • Unified interface: chat(messages, model, stream=True) is the only entry; internally it picks the platform HTTP implementation (URLSession on iOS, OkHttp on Android).
  • Weak-network retries: exponential backoff + jitter at the network layer, up to 3 retries; retryable on timeout / connection reset, not on 4xx.
  • Streaming rendering: push SSE chunks to the UI thread immediately, isolated via DispatchQueue.Main / MainScope to prevent UI stutter.
  • Battery budget: model calls are tagged opportunistic; in low-battery mode they auto-degrade to a cheaper model; background calls require charging + Wi-Fi.
  • Offline cache: responses for the same prompt land in SQLite; when offline, return the latest cache with a transparent "offline result" hint.
  • Token security: the user token lives in Keychain / Keystore, never in NSUserDefaults / SharedPreferences.

Code Example (Swift pseudocode)

class APIShareClient {
    let session = URLSession(configuration: .default)
    let cache = ResponseCache()  // SQLite-backed
    let token = Keychain.read("apshare_token")

    func chat(_ messages: [Message], model: String,
              onChunk: @escaping (String) -> Void) async throws {
        // 1. try cache
        if let cached = cache.get(messages, model) {
            onChunk(cached); return
        }
        // 2. streaming request with retry
        let req = try buildRequest(messages, model)
        for attempt in 0..<3 {
            do {
                let (bytes, _) = try await session.bytes(for: req)
                var buffer = ""
                for try await line in bytes.lines {
                    guard line.hasPrefix("data: ") else { continue }
                    let json = try JSON(line.dropFirst(6))
                    let delta = json.choices[0].delta.content
                    await MainActor.run { onChunk(delta) }
                    buffer += delta
                }
                cache.set(messages, model, buffer)
                return
            } catch let e as URLError where e.isRetryable {
                try await Task.sleep(nanoseconds: backoffNs(attempt))
                continue
            }
        }
        throw APIError.timeout
    }
}

App Update Granularity and SDK Compatibility

The biggest constraint on mobile SDKs is "users do not update": you ship a new SDK version fixing a bug, but 30% of users are still on last month's build. The SDK must make a backward compatibility promise: major version stays stable for 3 years, minor versions only add fields (never remove), and field semantics never change. Combined with a "config push" mechanism — the gateway pushes flags that control SDK behavior (disable a provider, tune retry counts) without requiring an app update.

Best Practices

  • Binary footprint: keep the SDK under 500KB so it does not bloat the app's main bundle.
  • Privacy manifest: iOS 17+ requires SDKs to declare PrivacyInfo.xcprivacy; a model-calling SDK must declare "no data collection."
  • A/B routing: the SDK supports experiment groups from apshare-config for easy model canarying.
  • Crash fallback: when JSON parsing of the gateway response fails, return a fixed fallback text instead of crashing.
  • Network tiering: route over Wi-Fi to high-quality models, over cellular to lightweight ones, saving battery without losing the experience.

Let apps share the unified gateway too: one interface, one set of retries, one cache strategy, consistent across platforms.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Cherry Studio Complete Guide: 300+ Models in One Desktop App — Local KB + MCP, Zero-Cost Unified CallingLobe Chat Complete Guide: Pluginized Web Unified Calling — Team KB & Visual Workflow, No-CodeOpen WebUI Complete Guide: Local Ollama + Cloud Free APIs in One Pool — Privacy-First Unified CallingPortkey AI Gateway Complete Guide: Enterprise Unified Calling for 250+ Models — Cache + Guardrails + ObservabilityLiteLLM Proxy Complete Guide: Python Unified Gateway for 100+ Models — OpenAI Compatible + Smart Routing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.