← Back to articles
Tutorials

Free Dialog System API Complete Tutorial: Build Context-Aware Conversational AI at Zero Cost (2026-10-10 Verified)

Free Dialog System API Complete Tutorial: Build Context-Aware Conversational AI at Zero Cost (2026-10-10 Verified)

One-line definition: Slot filling extracts predefined field values from natural language, converting unstructured conversations into machine-readable structured data.


1. What Is Dialog System

Slot filling is the core subtask of task-oriented dialog systems. When a user says "I want to book a flight from Beijing to Shanghai next Monday", the intent classification module identifies this as a "book_flight" intent, and the dialog system module must extract three required slots: departure_city=Beijing, destination_city=Shanghai, and travel_date=next_Monday. If any of these slots are missing or incorrect, the downstream flight search API cannot return meaningful results. This dependency chain is why dialog system errors are so costly in production — a single missing slot can cause the entire conversation to fail, requiring the dialog system to ask a clarification question and restart the slot collection process from scratch. When a user says "book me a flight tomorrow from Beijing to Shanghai," the system must automatically extract the departure city, destination city, travel date, and other fields from that sentence. This process is dialog system — mapping a piece of natural language to predefined structured fields.

Unlike named entity recognition, which only discovers entities, dialog system additionally maps entities to designated fields and handles constraint relationships between fields, such as ensuring the departure city differs from the destination. The output of dialog system is structured key-value pairs that can be directly fed into business logic layers for subsequent processing, including flight inventory queries, price calculations, and order generation. This capability to convert natural language into structured data is one of the key features that distinguish dialog systems from traditional search engines.


Task Goal Covered Here
Dialog System Extract predefined field values from dialog This article
Named Entity Recognition Discover entity types in text Free Named Entity Recognition (NER) API Complete Tutorial
Intent Classification Determine what the user wants Free Intent Classification API Complete Tutorial
Question Answering Answer user questions given context Free Question Answering (QA) API Complete Tutorial
Semantic Textual Similarity Measure semantic similarity between texts Free Semantic Textual Similarity (STS) API Complete Tutorial

Slot filling is often confused with NER, but they serve different purposes. NER answers the question "what entities exist in this text?" while dialog system answers "which fields does my business logic need, and what are their values?" A NER system might identify "Beijing" as a location and "tomorrow" as a date, but only dialog system knows that this particular dialog flow requires exactly three fields: origin, destination, and departure_date. Slot filling is therefore schema-driven, whereas NER is schema-free. This distinction matters when you design the data contract between your dialog system and its backend services. Imagine a travel booking dialog: the NER model extracts all location entities, but dialog system knows which of those are origin_city, which are destination_city, and which are irrelevant. A user might say "I want to fly from Beijing to Shanghai next Monday" and also mention "I previously visited Guangzhou" — NER extracts all three cities, but dialog system must assign only Beijing and Shanghai to the correct slots and ignore Guangzhou as historical context. This schema-driven filtering is what makes dialog system indispensable for production dialog systems. In a real customer service scenario, the user might say "I ordered product ABC on October 5th but received a different color" — NER extracts product=ABC, date=October 5th, and color=different, but dialog system must map these to order_id, expected_color, and received_color, then validate that the received_color differs from expected_color. Without this mapping and validation layer, the backend system would receive raw entities and have no idea which field corresponds to which business concept. This schema-driven filtering is what makes dialog system indispensable for production dialog systems. In real-world dialog systems, these five capabilities are typically chained together: intent classification first determines the user's goal, dialog system then extracts field values, and question answering handles follow-up queries. For the complete free NLP API catalog, visit our Free Keyword Extraction & Topic Modeling API Complete Tutorial page.


3. Free Channel Comparison

The radar chart reveals that spaCy local deployment scores perfectly across Chinese support, free quota, ease of use, response speed, and batch processing dimensions, making it the optimal choice for teams with zero budget. Azure AI Language stands alone in the custom slots dimension, making it ideal for enterprise scenarios requiring fine-grained field control. Hugging Face delivers balanced performance across ease of use and Chinese model richness, making it the preferred choice for rapid prototyping and proof-of-concept development. When selecting a channel, developers should also consider response latency, concurrency limits, model update frequency, and community activity levels. For high-throughput production deployments, the concurrency limit is often the binding constraint — if your peak traffic is 100 requests per second and your chosen channel only allows 10 concurrent requests, you will need to implement request queuing or switch to a different channel. The Hugging Face free tier typically allows 5 to 10 concurrent requests per account, while Google Cloud Natural Language allows up to 600 concurrent requests in the free tier — a significant difference for batch processing workloads., making it the optimal choice for teams with zero budget. Azure AI Language stands alone in the custom slots dimension, making it ideal for enterprise scenarios requiring fine-grained field control. Hugging Face delivers balanced performance across ease of use and Chinese model richness, making it the preferred choice for rapid prototyping and proof-of-concept development. When selecting a channel, developers should also consider response latency, concurrency limits, model update frequency, and community activity levels.


4. Channel Deep Dive

4.1 Hugging Face Inference API

The Hugging Face Inference API free tier enforces account-wide rate limits rather than per-model billing. Registration provides immediate access with no credit card required. For dialog system, the recommended approach is to use an NER model for entity extraction, then apply rule-based or lightweight models for slot assignment. The free tier typically allows 30 to 100 requests per minute depending on model size and current server load. When encountering 429 rate limit errors, implement exponential backoff with jitter to handle temporary throttling gracefully. The Hugging Face Inference API returns a standard HTTP 429 status code with a Retry-After header indicating how many seconds to wait before retrying. A robust implementation should parse this header and wait the specified duration before retrying, with a maximum retry limit of 3 to 5 attempts to avoid infinite loops. For batch processing workloads, consider using the Hugging Face Async Inference API, which allows you to submit a batch of requests and poll for results asynchronously — this pattern is more efficient than synchronous sequential processing for large batches. The Hugging Face Hub currently hosts over 200,000 NER-related models, many fine-tuned on Chinese datasets such as CMRC and DRCD, making it straightforward to find a model that matches your domain and language requirements. The free tier rate limit is account-wide, typically allowing 30 to 100 requests per minute depending on model size and current server load. When encountering 429 rate limit errors, implement exponential backoff with jitter to handle temporary throttling gracefully. For batch processing workloads, consider using the Hugging Face Inference Endpoints dedicated deployment option, which offers higher rate limits and more predictable latency at the cost of a paid plan — but for prototyping and low-volume production, the free tier is more than sufficient.

4.2 Google Cloud Natural Language API

The Google Cloud Natural Language API provides a permanent free quota of 5000 units per month with no credit card required. Each entity analysis request consumes approximately 1 to 2 units, allowing individual developers to process thousands of conversations monthly. The service supports over 20 languages, and its Chinese entity recognition accuracy ranks among the best in the free tier category. Beyond entity extraction, Google Cloud Natural Language also supports entity disambiguation and sentiment analysis, enabling simultaneous entity extraction and sentiment polarity detection — a capability particularly valuable for complaint identification in customer service scenarios. The API also offers syntax analysis and content classification features that can complement dialog system in more complex dialog systems. Content classification can be used to filter out irrelevant user utterances before they reach the dialog system module, reducing unnecessary API calls and improving overall system efficiency. The free tier includes 5,000 units per month, which is sufficient for approximately 2,500 to 5,000 entity analysis requests depending on document length — enough for a small-to-medium production workload without any cost.

4.3 Azure AI Language

The Azure AI Language F0 free tier provides 5000 text analysis requests per month without requiring a credit card. Azure's unique advantage lies in its custom slot template support — you can predefine fields such as departure city, destination, and travel date, and Azure automatically maps extracted entities to the corresponding slots. This capability is unmatched among the four free channels reviewed here, making it exceptionally valuable for dialog system developers. Azure also supports multilingual custom question answering projects where you can upload FAQ documents and the service automatically extracts question-answer pairs, making it particularly efficient for building enterprise knowledge base dialog systems. The free tier includes up to 3 projects with 100 documents each, sufficient for a medium-sized knowledge base. The Azure Language service also supports custom entity recognition projects where you define your own entity types and train the model on your annotated data. This is particularly powerful for domain-specific dialog system scenarios like healthcare, finance, or legal, where generic NER models perform poorly on specialized terminology. The free tier includes up to 5 custom entity types per project with 1,000 training documents, which is enough for most small-to-medium domain adaptation tasks.

4.4 spaCy Local Open Source

The Microsoft Azure AI Language service provides a custom question answering capability, but for dialog system specifically, its Language Understanding (LUIS) offering was the traditional choice. Microsoft has since integrated LUIS into Azure AI Language, where the conversational language understanding feature lets you define intents and entities (slots) through a web portal or API, then train and publish a model. The free tier (F0) is limited to 5,000 text records per month across language understanding, which is enough for prototyping and small production workloads. For developers who prefer to avoid external API dependencies entirely, spaCy offers a Chinese model zh_core_web_sm that is completely free, requires no network access, and imposes no usage limits. The model file is approximately 50 MB, loads extremely fast, and is suitable for embedded devices and offline scenarios. spaCy's pipeline architecture supports custom components, allowing you to add slot assignment, constraint validation, and other custom logic after NER without modifying the underlying model. This extensibility makes spaCy particularly attractive for teams with specific domain requirements that cannot be satisfied by off-the-shelf APIs. For teams with no GPU budget, the CPU-only spaCy inference is fast enough for sub-50-millisecond response times on standard server hardware. The zh_core_web_sm model file is only about 50 MB, so it loads quickly and runs comfortably on machines with 2 GB of RAM. For larger teams with GPU access, the larger transformer-based models can squeeze out an additional 10 to 15 percentage points of F1 score at the cost of longer inference times. The choice between DistilBERT and RoBERTa for dialog system is particularly interesting because dialog system benefits more from fine-grained token-level attention than from broad contextual understanding. RoBERTa's larger attention heads can capture subtle patterns like "next Monday" mapping to a date slot or "from Beijing" indicating a departure city — patterns that smaller models often miss. However, for most practical applications, the difference between 88% and 92% F1 is not worth the 5x increase in inference latency, and the rule-based fallback layer can bridge most of the remaining gap. The key insight is that you do not need the largest model — you need the right model for your specific slot types and language mix. For Chinese dialog system specifically, the zh_core_web_sm model achieves 85% F1 on standard benchmarks, which is sufficient for most practical applications. If you need higher accuracy, the zh_core_web_md and zh_core_web_lg variants offer incremental improvements at the cost of larger file sizes and slower inference. The spaCy training API allows you to fine-tune these models on your own annotated data, which is particularly valuable for domain-specific slot types like medical codes, product SKUs, or internal project names that are unlikely to appear in generic training data. For example, a medical dialog system might need to recognize drug names, dosages, and frequency intervals — none of which are covered by generic NER models. With spaCy, you can add a custom entity ruler that recognizes these domain-specific entities and maps them to your slot schema without retraining the entire pipeline. For more free NLP API comparisons, visit our Free Text Summarization page. The Azure Language service also supports custom entity recognition projects where you define your own entity types and train the model on your annotated data. This is particularly powerful for domain-specific dialog system scenarios like healthcare, finance, or legal, where generic NER models perform poorly on specialized terminology. The free tier includes up to 5 custom entity types per project with 1,000 training documents, which is enough for most small-to-medium domain adaptation tasks.


5. Code Example

The following Python example demonstrates the core logic for completing dialog system using the Hugging Face Inference API. The pattern is universal: call an NER model, map recognized entity types to your slot schema, then apply validation rules. The same structure works whether you call Google Cloud Natural Language, Azure AI Language, or a local spaCy model. For cost control and multi-model fallback strategies, see our Free API Cost and Quota Control in Practice page. In practice, the most robust dialog system pipelines combine rule-based extraction with LLM-based validation: first use NER or regex patterns to extract candidate slot values, then use an LLM to validate whether the extracted values are semantically correct in context. For example, if a user says "I want to fly from Beijing to Shanghai next Monday", the NER model might extract "Beijing" as a location, but only the LLM can confirm that Beijing is the departure city and Shanghai is the destination — not the other way around. This two-stage validation pattern dramatically reduces slot swapping errors. For cost control and multi-model fallback strategies, see our Free API Cost and Quota Control in Practice page.

import requests

API_URL = "https://huggingface.co/dslim/bert-base-NER"
headers = {"Authorization": "Bearer hf_your_token"}

def slot_fill(text):
    response = requests.post(API_URL, headers=headers, json={"inputs": text})
    entities = response.json()
    slots = {}
    for ent in entities:
        label = ent["entity_group"]
        value = text[ent["start"]:ent["end"]]
        slots[label] = value
    return slots

result = slot_fill("book me a flight tomorrow from Beijing to Shanghai")
print(result)

Note: The HF endpoint may be unreachable in certain network environments. Consider deploying the model yourself via the HF Hub page.


6. Dialog System vs NER: Key Differences

Our Free Named Entity Recognition (NER) API Complete Tutorial covers entity extraction in detail. Slot filling builds upon NER by adding two additional layers: field mapping, which maps NER output entity types to business-defined slot names, and constraint validation, which checks whether slot values satisfy business rules such as requiring dates to be in the future or ensuring the departure city differs from the destination city. These validation rules are typically defined as a formal schema that the dialog system module enforces before passing structured data to downstream business logic.


7. Building a Minimal Dialog System Pipeline

A functional dialog system pipeline requires only three steps: user input passes through NER entity extraction, then slot assignment, and finally produces structured output. Step one uses a Hugging Face NER model for entity extraction. Step two trains a lightweight classifier using rules or a small amount of annotated data to map entities to slots. Step three uses fuzzy matching methods from our Free Semantic Textual Similarity (STS) API Complete Tutorial to handle synonyms and aliases. In practice, we recommend starting with rule-based matching to quickly validate business logic, then gradually introducing machine learning models to improve accuracy. Rule-based matching offers transparent interpretability and low debugging costs, while machine learning models offer strong generalization and the ability to handle unseen expression patterns. A hybrid approach is often best: use ML for the initial extraction pass, then apply rules for validation and correction. For example, if the ML model extracts a date that is in the past for a booking intent, the rule layer can reject it and prompt the user for a valid date. This combination significantly reduces downstream errors without requiring expensive retraining cycles. In practice, we recommend starting with rule-based matching to quickly validate business logic, then gradually introducing machine learning models to improve accuracy. The rule layer should enforce constraints like required fields, value ranges, and type conversions. For example, a flight booking dialog might require the departure date to be in the future, the departure city to differ from the destination city, and the passenger count to be a positive integer. These rules are easy to write and debug, and they catch the most common user errors before they reach the backend. The ML layer should handle fuzzy matching, synonym resolution, and context-aware disambiguation. For example, if the user says "next Friday" without specifying a date, the ML layer should resolve this relative date expression to an absolute date based on the current date. Together they form a robust dialog system system that can handle the vast majority of real-world user inputs. The key to a successful hybrid system is to define clear boundaries between the rule layer and the ML layer, and to monitor their performance separately so you can improve each independently. Combining both approaches typically yields the best results.


8. Cost Comparison Across Channels

Below is a realistic cost projection for a dialog system pipeline handling 1,000 conversations per day, each requiring one NER call for entity extraction plus one slot assignment step. All channels remain within their free tiers at this volume.

Channel Monthly Call Volume Free Tier Capacity Cost
Hugging Face Inference API 60,000 Unlimited (rate-limited) $0
Google Cloud Natural Language 60,000 30,000 ($1/1,000 units excess) $0–$30/month
Azure AI Language (F0) 60,000 5,000 Requires S0 tier for excess
spaCy (local) Unlimited Unlimited $0 (hardware cost only)

For most small to medium projects, the Hugging Face or spaCy route keeps costs at exactly zero. The Google Cloud route becomes useful when you need entity sentiment analysis alongside dialog system, but the free tier is less generous than Hugging Face. Azure is best for teams already invested in the Microsoft ecosystem who want the conversational language understanding features. For cost control strategies, see our Free API Cost and Quota Control in Practice page.

9. FAQ

Q: What is the difference between dialog system and NER? A: NER only discovers entities. Slot filling additionally determines which field each entity belongs to and whether the field value is valid.

Q: Can free APIs handle multi-turn conversations? A: The free tiers of Hugging Face and Google Cloud are stateless single-call APIs. Multi-turn conversation context must be maintained by your own application logic.

Q: How accurate is Chinese dialog system? A: Hugging Face hosts numerous Chinese fine-tuned models that achieve near-commercial accuracy on Chinese entity recognition benchmarks. Models like uer/roberta-base-finetuned-chinaner-chinese and hfl/chinese-macbert-base typically score above 90 F1 on MSRA and OntoNotes datasets. For domain-specific scenarios such as healthcare or finance, further fine-tuning on domain corpora typically requires only a few thousand annotated examples to achieve significant improvements.

Q: Where should I start? Browse our Free API Cost and Quota Control in Practice to select a channel, and use our Free Text Summarization API Complete Tutorial registration link to experience the complete solution. We recommend starting with Hugging Face: register for an account, obtain an Access Token, and run through the first dialog system example in under five minutes. Once validated, migrate to other channels based on actual requirements. For teams with strict latency requirements, consider deploying a local spaCy model with custom entity rulers. For teams with complex schema validation, consider building a rule layer on top of any of the free NER APIs. The key insight is that dialog system is not a single API call — it is a pipeline that combines extraction, mapping, and validation. Start simple, measure honestly, and iterate based on real user queries. The first version of your dialog system system should use a single NER model with a basic rule layer. Measure its accuracy on a representative sample of real user inputs. Identify the most common failure modes — typically missing slots, wrong slot assignments, or values that fail validation — and address them one at a time. As your accuracy improves, you can add more sophisticated ML models for harder cases. The goal is not to achieve perfect accuracy from day one, but to build a system that improves continuously as you collect more data and refine your rules and models. For teams with no ML expertise, the rule-based approach alone can achieve 85–90% accuracy on well-defined dialog flows, which is sufficient for many production applications. We recommend starting with Hugging Face: register for an account, obtain an Access Token, and run through the first dialog system example in under five minutes. Once validated, migrate to other channels based on actual requirements.

More in this category

Free Slot Filling API Complete Tutorial: Extract Structured Fields from Conversations at Zero Cost (2026-10-10 Verified)Free Question Answering (QA) API Complete Tutorial: Give Your Text the Ability to Understand What Is Being Asked at Zero Cost (Verified 2026-10-09)Free Text Summarization API Complete Tutorial: Let LLMs Compress 1M-Word Documents into 100 WordsFree Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)Free Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.