← Back to articles
Detailed Usage

Using the Hugging Face Inference Client

⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification

Introduction

The Hugging Face Hub hosts hundreds of thousands of open-source models. With InferenceClient you can call most of them in a single line without deploying anything yourself. The free HF_INFERENCE endpoint is ideal for prototyping and lightweight tasks. This article demos text, streaming, cross-modal, and embedding usage.

架构图

flowchart LR A[pip install huggingface_hub] --> B[InferenceClient init] B --> C[Pick any model from HF Hub] C --> D[Text generation] C --> E[Image generation] C --> F[Audio transcription] D --> G[One-line call] E --> G F --> G

Installation

pip install -U huggingface_hub

Create a Read-scoped token at https://huggingface.co/settings/tokens:

export HF_TOKEN="hf_..."

Text Chat

import os
from huggingface_hub import InferenceClient

client = InferenceClient(token=os.environ["HF_TOKEN"])

answer = client.chat_completion(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Explain an inverted index in one sentence."},
    ],
    max_tokens=128,
)
print(answer.choices[0].message.content)

Streaming

for chunk in client.chat_completion(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "Write a haiku about autumn."}],
    stream=True,
):
    print(chunk.choices[0].delta.content or "", end="", flush=True)
print()

Cross-modal Calls

InferenceClient also wraps non-text tasks, with one method per task type:

# Text to image
img = client.text_to_image(
    "a cyberpunk cat, neon, ultra-detailed",
    model="black-forest-labs/FLUX.1-dev",
)
img.save("cat.png")

# Text to speech
with open("hello.wav", "wb") as f:
    f.write(client.text_to_speech(
        "Hello world", model="facebook/mms-tts-eng"
    ))

Text Embeddings

For RAG, you need to embed text into vectors:

vec = client.feature_extraction(
    "An inverted index is the core data structure of search engines.",
    model="intfloat/multilingual-e5-large",
)
print(vec.shape)  # (1, 1024)

Advanced: Let the Router Pick

Not sure which model to use? Let HF Router pick the best one for the task type:

# Without model=, the Router auto-selects the best model for the task
text = client.text_generation(
    "Once upon a time",
    model=None,
)

Dedicated Endpoints

The free HF_INFERENCE endpoint is strictly rate-limited. For production, upgrade to Dedicated Endpoints:

Task Routing Tips

Pick the right model per task: chat and light QA on meta-llama/Llama-3.1-8B-Instruct, reasoning and code on a 70B variant, Chinese workloads on Qwen/Qwen2.5-7B-Instruct, multimodal understanding on Qwen/Qwen2-VL-7B-Instruct. Each model card on the Hugging Face Hub lists the recommended task types — consult it during selection.

Troubleshooting

  • 429 rate limited: The free HF_INFERENCE endpoint is rate-limited. Slow down or upgrade to PRO ($9/month) for 20x quota.
  • model not loaded: Some large models are not on the free inference list — use a Dedicated Endpoint (paid).
  • Local inference: pip install transformers to load weights directly; small models run on CPU.
  • Gateway Timeout: Free endpoints can cold-start for ~30 seconds. Retry once.

InferenceClient is the fastest way to sample open-source models — ideal for the model-selection phase.

Best Practices

  • Reuse InferenceClient as singleton: do not new a client per request; the connection pool will be exhausted.
  • Set timeout to 60s: large model inference is slow; the default 30s often times out.
  • Expect cold start: HF free-tier models take 30-60s to load into GPU on the first call; subsequent calls are fast.
  • Separate clients for text vs image: text and image need different params (image_size, num_inference_steps); mixing them leads to errors.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →


Get Started: APIShare Free API Directory


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free AI Content Moderation API Guide 2026: Llama Guard 3 vs Perspective vs OpenAIFree OCR and Document Parsing API in PracticeIntegrating Free APIs into Your Local IDEConnecting Free Models to OpenCode in PracticeApplying for an OpenRouter API Key and Understanding Pricing

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.