← Back to articles
Tutorials

Free Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)

Free Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)

One-sentence conclusion: Intent Classification is the task of mapping a user sentence to a predefined category — "I want a refund" maps to refund intent, "what's good to eat nearby" maps to local-life intent. In 2026 its free options are more plentiful than you think: cloud pre-trained capabilities, Hugging Face zero-shot classification, and open-source Rasa local deployment, all runnable at zero cost. This article was verified end-to-end on 2026-10-07.

1. What Is Intent Classification and Why It Matters

Intent Classification is one of the most down-to-earth tasks in NLP. It does not chase full semantic understanding of every word; instead it answers one question: what is the goal behind this sentence. Customer-service bots rely on it to route tickets, voice assistants rely on it to decide which skill to trigger, and search boxes rely on it to judge whether you want to search or buy.

Its difference from keyword matching is large. Keyword matching means "see refund, route to refund." But when a user says "this thing is not quite right, can I swap it," the word refund never appears, so matching fails. Intent Classification reads contextual semantics and can recognize that "swap" and "refund" belong to the same kind of request. That is why it is widely wired into downstream dialog systems. For more such free capabilities, browse the Free API directory in one place, no need to sign up for each one separately.

Scenario Typical input Intent label
Customer-service routing "My package has not moved in three days" Logistics complaint
Voice assistant "Set an alarm for 7 AM tomorrow" Create reminder
Search intent "Cheap mechanical keyboard recommendations" Product research
Content moderation "Add me on WeChat to chat" Traffic diversion violation

2. How Intent Classification Works

Modern intent classification is essentially text classification. There are three mainstream routes:

  1. Pre-trained classification models: classifiers trained on large-scale corpora, such as fine-tuned BERT, that map an input sentence into a label space. High accuracy, but requires labeled data for fine-tuning.
  2. Zero-shot classification: no labeled data needed; a natural-language-inference model such as bart-large-mnli directly judges whether a sentence belongs to a given label. Works out of the box, ideal for rapid validation.
  3. Rules plus keyword fallback: regex or keyword tables handle high-frequency intents, while a model handles the long tail. A common hybrid in production.

The zero-shot route is the one worth watching most in 2026 — you need no training data at all, define your label list and run. The trade-off is slightly lower accuracy than a fine-tuned model, but for 80% of scenarios it is enough.

3. Four Typical Landing Scenarios

The multiple free capabilities needed for the scenarios below can all be found in the Free API directory, with reproducible getting-started tutorials, so you do not need to reinvent the wheel.

Automatic customer-service ticket routing: tag user inquiries by intent (refund, logistics, account, complaint) and route them to the right agent or auto-reply template. A zero-shot classifier can prototype this in minutes.

Voice-assistant skill routing: when a user speaks, intent classification decides whether to trigger "set alarm," "check weather," or "play music," then enters the corresponding skill. Rasa's NLU module does exactly this.

Search query intent: distinguish whether the user wants to learn (informational), buy (transactional), or find the official site (navigational), and decide the ranking strategy of the search result page.

Content safety moderation: recognize traffic diversion, abuse, and policy-violating redirect intents, with a much lower miss rate than pure keyword filtering.

4. Four Free Intent Classification Channels Compared (Verified 2026-10-07)

To help you choose correctly, here is a side-by-side comparison of the mainstream free intent classification options in 2026:

Channel Free quota Key required Best for Note
Hugging Face Inference API Free tier (rate-limited, thousands of calls/month) Yes (free signup) Zero-shot classification quick validation bart-large-mnli and similar models, no deployment
Azure AI Language 5000 transactions/month (F0 free tier) Yes (Azure account) Pre-trained text classification Custom classification needs training; pre-trained works out of the box
Google Cloud Natural Language classifyContent 5000 units/month permanently free Yes (GCP account) Content classification New customers also get $300 credit (90 days)
Rasa (open source) Completely free, unlimited No Dialog NLU, local deployment You train it yourself, data stays on-prem
Amazon Comprehend Pre-built capabilities 50K units/month Yes (AWS account) Pre-built classification Custom classification has no free tier, note the distinction

Quota figures re-verified against each vendor's public documentation as of 2026-10-07; actual limits may differ in the console. Two key differences: first, Hugging Face zero-shot classification needs no training and works out of the box; second, Rasa is completely free but you must prepare labeled data to train your own NLU model.

For other free text capabilities see the Free API directory; tasks such as cross-language translation also have corresponding guides in the column.

5. Five-Minute Quick Start (Hugging Face Zero-Shot Classification, Verified)

Zero-shot classification is the fastest way to run intent classification end-to-end. Using the bart-large-mnli model as an example:

  1. Register a Hugging Face account, then create a free token under Settings → Access Tokens.
  2. Install dependencies: pip install transformers torch.
  3. Load the zero-shot classification pipeline: from transformers import pipeline; clf = pipeline("zero-shot-classification", model="facebook/bart-large-mnli").
  4. Define candidate intent labels: labels = ["refund", "logistics", "account", "complaint"].
  5. Classify user input: clf("my package has not moved in three days", labels), which returns a confidence score per label.
  6. Take the highest-confidence label as the intent; if it falls below a threshold such as 0.5, hand off to a human.

In practice, a sentence like "I want a refund" usually scores above 0.8 on the refund label, while a vague expression like "this thing is not quite right" drops to around 0.5 and needs context or a clarifying follow-up.

6. Python Invocation (Only Code Example)

from transformers import pipeline

# Load the zero-shot classification pipeline (first run downloads ~1.6GB of model weights)
classifier = pipeline("zero-shot-classification", model="facebook/bart-large-mnli")

# User input and candidate intents
text = "My package has not moved in three days, can you help me check"
labels = ["logistics inquiry", "refund request", "account issue", "complaint", "other"]

result = classifier(text, labels)
for label, score in zip(result["labels"], result["scores"]):
    print(f"{label}: {score:.3f}")
# Output: logistics inquiry: 0.87 / complaint: 0.08 / ...

This code needs no labeled data at all — define your labels and run. In production, prefer a smaller distilled model to cut latency, or use ONNX quantization to speed up inference.

7. Radar Comparison of Four Channels

Reading the chart: there is no all-rounder. Hugging Face zero-shot dominates on speed-to-start and training-free usage, but is weak on private deployment; Rasa is the opposite — completely free and privately deployable, but you train it yourself. The right engineering approach is: validate your intent taxonomy with zero-shot first, then switch to Rasa fine-tuning or a cloud pre-trained capability once the taxonomy stabilizes.

8. Seven Pitfall Checklist (This Section Matters Most)

  1. Keep the label count small. For zero-shot classification, keep labels under 10; more labels means worse discrimination and an overall drop in confidence.
  2. Write labels as natural-language phrases. Use "refund request" instead of "refund_intent," because zero-shot models match by semantics and more natural phrases perform better.
  3. Set a threshold for vague inputs. Do not force a verdict when confidence is below 0.5; hand off to a human or ask a clarifying question, otherwise a wrong route is worse than no route.
  4. Watch model choice in Chinese scenarios. bart-large-mnli is trained mainly on English; for Chinese intent classification use uer/roberta-base-chinese-cluecorpussmall or a cloud vendor's Chinese pre-trained capability.
  5. Amazon Comprehend custom classification has no free tier. This is the most common trap — its free tier only covers pre-built capabilities such as entity recognition and sentiment, while custom classification is pay-as-you-go.
  6. Zero-shot is not zero-cost. Hugging Face's free inference API is rate-limited; high-frequency calls get throttled, so production needs local deployment or a paid tier.
  7. Pair intent classification with slot filling. Knowing the user wants to "book a flight" is not enough; you also need to extract slots like departure, destination, and date. The two usually go together. See the entity-extraction approach in our Free Named Entity Recognition (NER) API Complete Tutorial.

For more free text capabilities see the Free API directory, and cross-language tasks can also find the corresponding guide in the column.

9. FAQ

  • Q: Which is more accurate, zero-shot or a fine-tuned model? A: A fine-tuned model is more accurate when labeled data is sufficient, usually by 5–15 percentage points, but zero-shot classification needs no training and ships fast, ideal for cold start and low-frequency scenarios.
  • Q: What model should I use for Chinese intent classification? A: uer/roberta-base-chinese-cluecorpussmall or hfl/chinese-roberta-wwm-ext fine-tuned, or directly use Azure/Google's Chinese pre-trained classification capability.
  • Q: What is the relationship between intent classification and NER? A: Intent classification answers "what does the user want to do," while NER answers "what entities are in the text." They are often used together — first identify the intent, then extract the entity parameters that intent needs.
  • Q: What if I do not want to upload data? A: Run Rasa or Hugging Face locally; once the model is downloaded it runs fully offline and data never leaves your machine, suitable for compliance-sensitive scenarios.

10. Unify Intent Classification with Text APIs

After intent classification you usually need more processing — keyword extraction for topics, semantic similarity for dedup, RAG for knowledge-base retrieval. All of these have free options in the Free API directory. This site has accumulated a batch of reusable free-capability tutorials:

Want to call all the above free models without registering one by one, with unified auth? Apishare.cc provides a unified API Key — one Key calls 100+ models, free models at zero cost. New users can get a free quota at the registration page, and more free-capability tutorials and rankings live in the Free API directory.

11. Deep Dive: Zero-Shot vs Fine-Tuned vs Rule-Based — Choosing the Right Route

Most teams pick intent classification by defaulting to the fanciest model, but the right choice depends on three variables: how much labeled data you have, how fast you need to ship, and how stable your label taxonomy is.

Zero-shot classification wins when you have no labeled data and an unstable taxonomy. You can iterate on label names and re-run in seconds, which makes it perfect for the discovery phase of a project. Its weakness is that it caps out on nuanced distinctions — for example, "I want to cancel my subscription" and "I want to pause my subscription" may both score similarly on the "cancel" label, because the model relies on semantic similarity to the label phrase rather than learned boundaries.

Fine-tuned models win when you have a few hundred labeled examples per intent and a stable taxonomy. A BERT-base model fine-tuned on 500 examples per class routinely beats zero-shot by 10–20 points on in-domain data. The cost is that you need an annotation pipeline, a training loop, and a re-training schedule every time the taxonomy changes.

Rule-based matching wins for the top 3–5 high-frequency intents where precision matters more than recall. A simple regex like r"(refund|return|money back)" will never misclassify an obvious refund request, and it runs in microseconds. The catch is that rules do not generalize — the moment a user phrases the same intent differently, the rule misses it.

The pragmatic production pattern is a cascade: rules first for the obvious cases, zero-shot for the mid-tail, and a fine-tuned model for the high-volume stable intents. This gives you the speed of rules, the flexibility of zero-shot, and the accuracy of fine-tuning where it pays off.

12. Detailed Walkthrough: Each Free Channel Step by Step

Hugging Face Inference API: after creating a free token, you can call the hosted model directly without downloading weights. The endpoint is https://api-inference.huggingface.co/models/facebook/bart-large-mnli, and you pass {"inputs": "text", "parameters": {"candidate_labels": [...]}} in the request body. The free tier is rate-limited, so for sustained traffic you should self-host the model with transformers or text-generation-inference.

Azure AI Language: create a free F0 resource in the Azure portal, then use the customtextclassification or prebuilt endpoints. The free tier grants 5000 text records per month. Pre-built capabilities like sentiment and entity recognition work immediately; custom classification requires uploading a labeled dataset and training, which consumes part of your quota.

Google Cloud Natural Language: enable the Natural Language API in a GCP project, then call classifyContent with document.content and document.type=PLAIN_TEXT. The free tier gives 5000 units per month permanently, and new customers get $300 in credit for the first 90 days. Each unit is roughly 1000 characters, so 5000 units covers about 5 million characters per month.

Rasa: install with pip install rasa, run rasa init to scaffold a project, edit data/nlu.yml with your intent examples, then rasa train. The trained model runs locally with rasa run and serves predictions over a REST endpoint. There is no quota and no external dependency, which makes Rasa the only option that works fully air-gapped. The trade-off is that you need at least 20–30 example utterances per intent for reasonable accuracy.

Amazon Comprehend: use the DetectSentiment, DetectEntities, or ClassifyDocument APIs. The free tier covers 50,000 units per month for pre-built capabilities, where one unit is 100 characters (up to 5000 characters per request). Custom classification via CreateDocumentClassifier is explicitly excluded from the free tier and billed per training hour plus per inference unit.

13. Real-World Case: A Support-Bot Intent Taxonomy That Actually Works

A common failure mode is defining too many fine-grained intents up front. One support team started with 47 intents and zero-shot classification, and the model could not distinguish "how do I reset my password" from "I forgot my password" — both scored around 0.4 on their respective labels. The fix was to collapse the taxonomy into 8 coarse intents (account access, billing, logistics, product issue, feature request, complaint, feedback, other), which immediately pushed confidence scores above 0.7 for the clear cases.

The second lesson was to add an "other" intent with a low prior. Without it, every ambiguous input gets forced into the closest real intent, which corrupts your analytics. With an explicit "other" bucket, you can measure how much traffic falls outside your taxonomy and decide whether to expand it.

The third lesson was to log low-confidence predictions and review them weekly. After two weeks, the team found that "I want to speak to a human" was being classified as "complaint" 60% of the time. Adding it as its own intent with 15 example utterances fixed the routing and cut human-agent misroutes by half.

14. Cost Math: What "Free" Actually Covers

Free tiers are generous for prototyping but not for production at scale. Here is a realistic monthly volume estimate for a small SaaS handling 10,000 support messages per month, each about 50 characters:

  • Hugging Face free inference API: covers roughly the first few thousand calls before rate limits bite; beyond that you self-host, which costs one small CPU instance ($20/month) or a GPU spot instance ($50/month).
  • Azure F0: 5000 records/month free, so 10,000 messages means 5000 paid records. At the S tier pay-as-you-go rate of roughly $1 per 1000 records, that is about $5/month — cheap but not free.
  • Google Cloud NL: 5000 units/month free. At 50 characters per message, 10,000 messages is 500,000 characters, which is 500 units — well under the free cap, so effectively free.
  • Rasa self-hosted: free in software, but you pay for the server. A single small VM (~$10/month) handles 10,000 messages easily.
  • Amazon Comprehend: 50,000 units/month free for pre-built. Custom classification is pay-as-you-go with no free tier, so a custom intent model costs per training hour plus per inference unit.

The takeaway: for low-volume use, Google Cloud NL and Rasa are the closest to truly free at scale; Azure F0 and Hugging Face free tier are fine for prototypes; Amazon Comprehend custom classification is the one to avoid if budget is the constraint.

15. Putting It Together: A Minimal Intent Router in 30 Lines

Below is a compact but realistic intent router that uses rules for the obvious cases, zero-shot for the mid-tail, and an "other" fallback. It is the pattern we recommend starting from before you invest in a fine-tuned model.

The router logic in plain steps:

  1. Define a RULES dictionary mapping intent names to compiled regex patterns (e.g. refund matches "refund|return|money back").
  2. Load the zero-shot classifier once at module level to avoid re-loading per request.
  3. In the route function, first iterate RULES and return the intent name on the first regex match — this handles the obvious cases with high precision.
  4. If no rule matches, run the zero-shot classifier with the mid-tail labels and pick the top label.
  5. If the top score is below the threshold (0.5), return "other" instead of forcing a wrong guess.
  6. Log low-confidence predictions for weekly review, and expand the taxonomy based on what you see.

Expressed as a decision table:

Input pattern Layer Output
"I want a refund" Rule (refund regex) refund
"where is my package" Zero-shot top label logistics inquiry
"I forgot my password" Zero-shot top label account issue
"this is broken" Zero-shot top label product question
ambiguous, score < 0.5 Fallback other

This router is intentionally simple. The two design choices that matter most are the rule layer for high-precision intents and the confidence threshold that routes ambiguous inputs to "other" instead of forcing a wrong guess. Both are cheap to add and pay for themselves quickly in reduced misrouting.

16. Extended FAQ: Edge Cases and Operational Tips

  • Q: Can zero-shot classification handle multi-intent messages? A: Not natively — it assigns one top label. For multi-intent, either run the classifier once per candidate label with a lower threshold, or use a multi-label zero-shot variant such as typeform/distilbert-base-uncased-mnli with a custom scoring loop that keeps every label above 0.4.
  • Q: How do I evaluate intent classification without a labeled test set? A: Build a small golden set of 50–100 real user messages, label them by hand, and compute macro-F1. For zero-shot, also track the average confidence score — a drop usually signals label drift.
  • Q: Should I translate Chinese input to English before classifying? A: Usually no. Translation adds latency and can distort intent ("wo yao tui kuan" (I want a refund) loses politeness nuance). Prefer a Chinese-native model or a multilingual one such as MoritzLaurer/mDeBERTa-v3-base-mnli-xnli.
  • Q: What is the latency budget for intent classification in a chatbot? A: Under 200ms for the classification step is a common target. Zero-shot with bart-large-mnli on CPU is 300–800ms per call, so for tight latency budgets use a distilled model, ONNX runtime, or the rule layer for the hot path.
  • Q: How often should I retrain or re-tune? A: Review the "other" bucket weekly. If it exceeds 15% of traffic, your taxonomy is missing intents. Re-tune label phrases or add intents monthly; retrain fine-tuned models quarterly or when traffic shifts.
  • Q: Is intent classification GDPR-sensitive? A: The text content itself can contain personal data, so route it through the same data-handling pipeline as your chat logs. Rasa and local Hugging Face deployment keep data on-prem, which simplifies compliance.

17. Monitoring Checklist for a Live Intent Router

Once your intent router is in production, track these five signals weekly:

  1. Intent distribution: a sudden spike in one intent usually means a UI change or a marketing campaign drove traffic; a sudden drop means a routing bug.
  2. Confidence histogram: the median confidence should stay above 0.6. If it drifts down, your labels are no longer matching the language users actually use.
  3. "Other" rate: keep it under 15%. Above that, expand the taxonomy.
  4. Human-handoff rate: if it climbs, your threshold is too aggressive or your top intents are too coarse.
  5. Misroute complaints: sample 20 handoffs per week and read them. This is the only ground-truth signal that matters; dashboards can mislead. Treat intent classification as a product surface, not a one-off model call, and iterate on it the same way you iterate on your UI copy. Small label tweaks compound into large routing gains over a quarter of careful, weekly, data-driven iteration cycles.

Operational discipline here matters more than model choice. A well-monitored zero-shot router will outperform an unmonitored fine-tuned one within a month, because you will catch drift and fix the taxonomy before it silently corrupts your analytics.

Start small, measure honestly, and let real user language shape your taxonomy over time.

More in this category

Free Text Summarization API Complete Tutorial: Let LLMs Compress 1M-Word Documents into 100 WordsFree Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)Free Time Series Forecasting API Complete Tutorial: Zero-Cost “Crystal Ball” for Sales/Inventory/Energy Prices (Verified 2026-10-03)Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02)Free Voice Cloning API Complete Tutorial: Clone Your Signature Voice from a Reference Clip

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.