⚠️ Pending Update · 2026-08-29 Verification · Content may be outdated, please refer to official docs Updated: 2026-08-29 · Status: Pending Verification
Introduction
The Hugging Face Hub hosts hundreds of thousands of open-source models. With InferenceClient you can call most of them in a single line without deploying anything yourself. The free HF_INFERENCE endpoint is ideal for prototyping and lightweight tasks. This article demos text, streaming, cross-modal, and embedding usage.
架构图
Installation
pip install -U huggingface_hub
Create a Read-scoped token at https://huggingface.co/settings/tokens:
export HF_TOKEN="hf_..."
Text Chat
import os
from huggingface_hub import InferenceClient
client = InferenceClient(token=os.environ["HF_TOKEN"])
answer = client.chat_completion(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain an inverted index in one sentence."},
],
max_tokens=128,
)
print(answer.choices[0].message.content)
Streaming
for chunk in client.chat_completion(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "Write a haiku about autumn."}],
stream=True,
):
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()
Cross-modal Calls
InferenceClient also wraps non-text tasks, with one method per task type:
# Text to image
img = client.text_to_image(
"a cyberpunk cat, neon, ultra-detailed",
model="black-forest-labs/FLUX.1-dev",
)
img.save("cat.png")
# Text to speech
with open("hello.wav", "wb") as f:
f.write(client.text_to_speech(
"Hello world", model="facebook/mms-tts-eng"
))
Text Embeddings
For RAG, you need to embed text into vectors:
vec = client.feature_extraction(
"An inverted index is the core data structure of search engines.",
model="intfloat/multilingual-e5-large",
)
print(vec.shape) # (1, 1024)
Advanced: Let the Router Pick
Not sure which model to use? Let HF Router pick the best one for the task type:
# Without model=, the Router auto-selects the best model for the task
text = client.text_generation(
"Once upon a time",
model=None,
)
Dedicated Endpoints
The free HF_INFERENCE endpoint is strictly rate-limited. For production, upgrade to Dedicated Endpoints:
- Create a dedicated GPU instance at https://ui.endpoints.huggingface.co
- Billed hourly, starting around $0.06/h (T4)
- Fully dedicated resources, no concurrent contention
Task Routing Tips
Pick the right model per task: chat and light QA on meta-llama/Llama-3.1-8B-Instruct, reasoning and code on a 70B variant, Chinese workloads on Qwen/Qwen2.5-7B-Instruct, multimodal understanding on Qwen/Qwen2-VL-7B-Instruct. Each model card on the Hugging Face Hub lists the recommended task types — consult it during selection.
Troubleshooting
429 rate limited: The freeHF_INFERENCEendpoint is rate-limited. Slow down or upgrade to PRO ($9/month) for 20x quota.model not loaded: Some large models are not on the free inference list — use a Dedicated Endpoint (paid).- Local inference:
pip install transformersto load weights directly; small models run on CPU. Gateway Timeout: Free endpoints can cold-start for ~30 seconds. Retry once.
InferenceClient is the fastest way to sample open-source models — ideal for the model-selection phase.
Best Practices
- Reuse InferenceClient as singleton: do not new a client per request; the connection pool will be exhausted.
- Set timeout to 60s: large model inference is slow; the default 30s often times out.
- Expect cold start: HF free-tier models take 30-60s to load into GPU on the first call; subsequent calls are fast.
- Separate clients for text vs image: text and image need different params (image_size, num_inference_steps); mixing them leads to errors.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Get Started: APIShare Free API Directory
- 🆓 Claim your free credits:Register on APIShare · Sign in to console
- 🔍 Browse every free API and live ranking:APIShare Free API Directory
- 📊 See the leaderboard:Free LLM API Rankings
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key