Free Slot Filling API Complete Tutorial: Extract Structured Fields from Conversations at Zero Cost (2026-10-10 Verified)
One-line definition: Slot filling extracts predefined field values from natural language, converting unstructured conversations into machine-readable structured data.
1. What Is Slot Filling
Slot filling is the core subtask of task-oriented dialog systems. When a user says "I want to book a flight from Beijing to Shanghai next Monday", the intent classification module identifies this as a "book_flight" intent, and the slot filling module must extract three required slots: departure_city=Beijing, destination_city=Shanghai, and travel_date=next_Monday. If any of these slots are missing or incorrect, the downstream flight search API cannot return meaningful results. This dependency chain is why slot filling errors are so costly in production — a single missing slot can cause the entire conversation to fail, requiring the dialog system to ask a clarification question and restart the slot collection process from scratch. When a user says "book me a flight tomorrow from Beijing to Shanghai," the system must automatically extract the departure city, destination city, travel date, and other fields from that sentence. This process is slot filling — mapping a piece of natural language to predefined structured fields.
Unlike named entity recognition, which only discovers entities, slot filling additionally maps entities to designated fields and handles constraint relationships between fields, such as ensuring the departure city differs from the destination. The output of slot filling is structured key-value pairs that can be directly fed into business logic layers for subsequent processing, including flight inventory queries, price calculations, and order generation. This capability to convert natural language into structured data is one of the key features that distinguish dialog systems from traditional search engines.
2. Slot Filling vs Related NLP Tasks
| Task | Goal | Covered Here |
|---|---|---|
| Slot Filling | Extract predefined field values from dialog | This article |
| Named Entity Recognition | Discover entity types in text | Free Named Entity Recognition (NER) API Complete Tutorial |
| Intent Classification | Determine what the user wants | Free Intent Classification API Complete Tutorial |
| Question Answering | Answer user questions given context | Free Question Answering (QA) API Complete Tutorial |
| Semantic Textual Similarity | Measure semantic similarity between texts | Free Semantic Textual Similarity (STS) API Complete Tutorial |
Slot filling is often confused with NER, but they serve different purposes. NER answers the question "what entities exist in this text?" while slot filling answers "which fields does my business logic need, and what are their values?" A NER system might identify "Beijing" as a location and "tomorrow" as a date, but only slot filling knows that this particular dialog flow requires exactly three fields: origin, destination, and departure_date. Slot filling is therefore schema-driven, whereas NER is schema-free. This distinction matters when you design the data contract between your dialog system and its backend services. Imagine a travel booking dialog: the NER model extracts all location entities, but slot filling knows which of those are origin_city, which are destination_city, and which are irrelevant. A user might say "I want to fly from Beijing to Shanghai next Monday" and also mention "I previously visited Guangzhou" — NER extracts all three cities, but slot filling must assign only Beijing and Shanghai to the correct slots and ignore Guangzhou as historical context. This schema-driven filtering is what makes slot filling indispensable for production dialog systems. In a real customer service scenario, the user might say "I ordered product ABC on October 5th but received a different color" — NER extracts product=ABC, date=October 5th, and color=different, but slot filling must map these to order_id, expected_color, and received_color, then validate that the received_color differs from expected_color. Without this mapping and validation layer, the backend system would receive raw entities and have no idea which field corresponds to which business concept. This schema-driven filtering is what makes slot filling indispensable for production dialog systems. In real-world dialog systems, these five capabilities are typically chained together: intent classification first determines the user's goal, slot filling then extracts field values, and question answering handles follow-up queries. For the complete free NLP API catalog, visit our Free Keyword Extraction & Topic Modeling API Complete Tutorial page.
3. Free Channel Comparison
The radar chart reveals that spaCy local deployment scores perfectly across Chinese support, free quota, ease of use, response speed, and batch processing dimensions, making it the optimal choice for teams with zero budget. Azure AI Language stands alone in the custom slots dimension, making it ideal for enterprise scenarios requiring fine-grained field control. Hugging Face delivers balanced performance across ease of use and Chinese model richness, making it the preferred choice for rapid prototyping and proof-of-concept development. When selecting a channel, developers should also consider response latency, concurrency limits, model update frequency, and community activity levels. For high-throughput production deployments, the concurrency limit is often the binding constraint — if your peak traffic is 100 requests per second and your chosen channel only allows 10 concurrent requests, you will need to implement request queuing or switch to a different channel. The Hugging Face free tier typically allows 5 to 10 concurrent requests per account, while Google Cloud Natural Language allows up to 600 concurrent requests in the free tier — a significant difference for batch processing workloads., making it the optimal choice for teams with zero budget. Azure AI Language stands alone in the custom slots dimension, making it ideal for enterprise scenarios requiring fine-grained field control. Hugging Face delivers balanced performance across ease of use and Chinese model richness, making it the preferred choice for rapid prototyping and proof-of-concept development. When selecting a channel, developers should also consider response latency, concurrency limits, model update frequency, and community activity levels.
4. Channel Deep Dive
4.1 Hugging Face Inference API
The Hugging Face Inference API free tier enforces account-wide rate limits rather than per-model billing. Registration provides immediate access with no credit card required. For slot filling, the recommended approach is to use an NER model for entity extraction, then apply rule-based or lightweight models for slot assignment. The free tier typically allows 30 to 100 requests per minute depending on model size and current server load. When encountering 429 rate limit errors, implement exponential backoff with jitter to handle temporary throttling gracefully. The Hugging Face Inference API returns a standard HTTP 429 status code with a Retry-After header indicating how many seconds to wait before retrying. A robust implementation should parse this header and wait the specified duration before retrying, with a maximum retry limit of 3 to 5 attempts to avoid infinite loops. For batch processing workloads, consider using the Hugging Face Async Inference API, which allows you to submit a batch of requests and poll for results asynchronously — this pattern is more efficient than synchronous sequential processing for large batches. The Hugging Face Hub currently hosts over 200,000 NER-related models, many fine-tuned on Chinese datasets such as CMRC and DRCD, making it straightforward to find a model that matches your domain and language requirements. The free tier rate limit is account-wide, typically allowing 30 to 100 requests per minute depending on model size and current server load. When encountering 429 rate limit errors, implement exponential backoff with jitter to handle temporary throttling gracefully. For batch processing workloads, consider using the Hugging Face Inference Endpoints dedicated deployment option, which offers higher rate limits and more predictable latency at the cost of a paid plan — but for prototyping and low-volume production, the free tier is more than sufficient.
4.2 Google Cloud Natural Language API
The Google Cloud Natural Language API provides a permanent free quota of 5000 units per month with no credit card required. Each entity analysis request consumes approximately 1 to 2 units, allowing individual developers to process thousands of conversations monthly. The service supports over 20 languages, and its Chinese entity recognition accuracy ranks among the best in the free tier category. Beyond entity extraction, Google Cloud Natural Language also supports entity disambiguation and sentiment analysis, enabling simultaneous entity extraction and sentiment polarity detection — a capability particularly valuable for complaint identification in customer service scenarios. The API also offers syntax analysis and content classification features that can complement slot filling in more complex dialog systems. Content classification can be used to filter out irrelevant user utterances before they reach the slot filling module, reducing unnecessary API calls and improving overall system efficiency. The free tier includes 5,000 units per month, which is sufficient for approximately 2,500 to 5,000 entity analysis requests depending on document length — enough for a small-to-medium production workload without any cost.
4.3 Azure AI Language
The Azure AI Language F0 free tier provides 5000 text analysis requests per month without requiring a credit card. Azure's unique advantage lies in its custom slot template support — you can predefine fields such as departure city, destination, and travel date, and Azure automatically maps extracted entities to the corresponding slots. This capability is unmatched among the four free channels reviewed here, making it exceptionally valuable for dialog system developers. Azure also supports multilingual custom question answering projects where you can upload FAQ documents and the service automatically extracts question-answer pairs, making it particularly efficient for building enterprise knowledge base dialog systems. The free tier includes up to 3 projects with 100 documents each, sufficient for a medium-sized knowledge base. The Azure Language service also supports custom entity recognition projects where you define your own entity types and train the model on your annotated data. This is particularly powerful for domain-specific slot filling scenarios like healthcare, finance, or legal, where generic NER models perform poorly on specialized terminology. The free tier includes up to 5 custom entity types per project with 1,000 training documents, which is enough for most small-to-medium domain adaptation tasks.
4.4 spaCy Local Open Source
The Microsoft Azure AI Language service provides a custom question answering capability, but for slot filling specifically, its Language Understanding (LUIS) offering was the traditional choice. Microsoft has since integrated LUIS into Azure AI Language, where the conversational language understanding feature lets you define intents and entities (slots) through a web portal or API, then train and publish a model. The free tier (F0) is limited to 5,000 text records per month across language understanding, which is enough for prototyping and small production workloads. For developers who prefer to avoid external API dependencies entirely, spaCy offers a Chinese model zh_core_web_sm that is completely free, requires no network access, and imposes no usage limits. The model file is approximately 50 MB, loads extremely fast, and is suitable for embedded devices and offline scenarios. spaCy's pipeline architecture supports custom components, allowing you to add slot assignment, constraint validation, and other custom logic after NER without modifying the underlying model. This extensibility makes spaCy particularly attractive for teams with specific domain requirements that cannot be satisfied by off-the-shelf APIs. For teams with no GPU budget, the CPU-only spaCy inference is fast enough for sub-50-millisecond response times on standard server hardware. The zh_core_web_sm model file is only about 50 MB, so it loads quickly and runs comfortably on machines with 2 GB of RAM. For larger teams with GPU access, the larger transformer-based models can squeeze out an additional 10 to 15 percentage points of F1 score at the cost of longer inference times. The choice between DistilBERT and RoBERTa for slot filling is particularly interesting because slot filling benefits more from fine-grained token-level attention than from broad contextual understanding. RoBERTa's larger attention heads can capture subtle patterns like "next Monday" mapping to a date slot or "from Beijing" indicating a departure city — patterns that smaller models often miss. However, for most practical applications, the difference between 88% and 92% F1 is not worth the 5x increase in inference latency, and the rule-based fallback layer can bridge most of the remaining gap. The key insight is that you do not need the largest model — you need the right model for your specific slot types and language mix. For Chinese slot filling specifically, the zh_core_web_sm model achieves 85% F1 on standard benchmarks, which is sufficient for most practical applications. If you need higher accuracy, the zh_core_web_md and zh_core_web_lg variants offer incremental improvements at the cost of larger file sizes and slower inference. The spaCy training API allows you to fine-tune these models on your own annotated data, which is particularly valuable for domain-specific slot types like medical codes, product SKUs, or internal project names that are unlikely to appear in generic training data. For example, a medical dialog system might need to recognize drug names, dosages, and frequency intervals — none of which are covered by generic NER models. With spaCy, you can add a custom entity ruler that recognizes these domain-specific entities and maps them to your slot schema without retraining the entire pipeline. For more free NLP API comparisons, visit our Free Text Summarization page. The Azure Language service also supports custom entity recognition projects where you define your own entity types and train the model on your annotated data. This is particularly powerful for domain-specific slot filling scenarios like healthcare, finance, or legal, where generic NER models perform poorly on specialized terminology. The free tier includes up to 5 custom entity types per project with 1,000 training documents, which is enough for most small-to-medium domain adaptation tasks.
5. Code Example
The following Python example demonstrates the core logic for completing slot filling using the Hugging Face Inference API. The pattern is universal: call an NER model, map recognized entity types to your slot schema, then apply validation rules. The same structure works whether you call Google Cloud Natural Language, Azure AI Language, or a local spaCy model. For cost control and multi-model fallback strategies, see our Free API Cost and Quota Control in Practice page. In practice, the most robust slot filling pipelines combine rule-based extraction with LLM-based validation: first use NER or regex patterns to extract candidate slot values, then use an LLM to validate whether the extracted values are semantically correct in context. For example, if a user says "I want to fly from Beijing to Shanghai next Monday", the NER model might extract "Beijing" as a location, but only the LLM can confirm that Beijing is the departure city and Shanghai is the destination — not the other way around. This two-stage validation pattern dramatically reduces slot swapping errors. For cost control and multi-model fallback strategies, see our Free API Cost and Quota Control in Practice page.
import requests
API_URL = "https://huggingface.co/dslim/bert-base-NER"
headers = {"Authorization": "Bearer hf_your_token"}
def slot_fill(text):
response = requests.post(API_URL, headers=headers, json={"inputs": text})
entities = response.json()
slots = {}
for ent in entities:
label = ent["entity_group"]
value = text[ent["start"]:ent["end"]]
slots[label] = value
return slots
result = slot_fill("book me a flight tomorrow from Beijing to Shanghai")
print(result)
Note: The HF endpoint may be unreachable in certain network environments. Consider deploying the model yourself via the HF Hub page.
6. Slot Filling vs NER: Key Differences
Our Free Named Entity Recognition (NER) API Complete Tutorial covers entity extraction in detail. Slot filling builds upon NER by adding two additional layers: field mapping, which maps NER output entity types to business-defined slot names, and constraint validation, which checks whether slot values satisfy business rules such as requiring dates to be in the future or ensuring the departure city differs from the destination city. These validation rules are typically defined as a formal schema that the slot filling module enforces before passing structured data to downstream business logic.
7. Building a Minimal Slot Filling Pipeline
A functional slot filling pipeline requires only three steps: user input passes through NER entity extraction, then slot assignment, and finally produces structured output. Step one uses a Hugging Face NER model for entity extraction. Step two trains a lightweight classifier using rules or a small amount of annotated data to map entities to slots. Step three uses fuzzy matching methods from our Free Semantic Textual Similarity (STS) API Complete Tutorial to handle synonyms and aliases. In practice, we recommend starting with rule-based matching to quickly validate business logic, then gradually introducing machine learning models to improve accuracy. Rule-based matching offers transparent interpretability and low debugging costs, while machine learning models offer strong generalization and the ability to handle unseen expression patterns. A hybrid approach is often best: use ML for the initial extraction pass, then apply rules for validation and correction. For example, if the ML model extracts a date that is in the past for a booking intent, the rule layer can reject it and prompt the user for a valid date. This combination significantly reduces downstream errors without requiring expensive retraining cycles. In practice, we recommend starting with rule-based matching to quickly validate business logic, then gradually introducing machine learning models to improve accuracy. The rule layer should enforce constraints like required fields, value ranges, and type conversions. For example, a flight booking dialog might require the departure date to be in the future, the departure city to differ from the destination city, and the passenger count to be a positive integer. These rules are easy to write and debug, and they catch the most common user errors before they reach the backend. The ML layer should handle fuzzy matching, synonym resolution, and context-aware disambiguation. For example, if the user says "next Friday" without specifying a date, the ML layer should resolve this relative date expression to an absolute date based on the current date. Together they form a robust slot filling system that can handle the vast majority of real-world user inputs. The key to a successful hybrid system is to define clear boundaries between the rule layer and the ML layer, and to monitor their performance separately so you can improve each independently. Combining both approaches typically yields the best results.
8. Cost Comparison Across Channels
Below is a realistic cost projection for a slot filling pipeline handling 1,000 conversations per day, each requiring one NER call for entity extraction plus one slot assignment step. All channels remain within their free tiers at this volume.
| Channel | Monthly Call Volume | Free Tier Capacity | Cost |
|---|---|---|---|
| Hugging Face Inference API | 60,000 | Unlimited (rate-limited) | $0 |
| Google Cloud Natural Language | 60,000 | 30,000 ($1/1,000 units excess) | $0–$30/month |
| Azure AI Language (F0) | 60,000 | 5,000 | Requires S0 tier for excess |
| spaCy (local) | Unlimited | Unlimited | $0 (hardware cost only) |
For most small to medium projects, the Hugging Face or spaCy route keeps costs at exactly zero. The Google Cloud route becomes useful when you need entity sentiment analysis alongside slot filling, but the free tier is less generous than Hugging Face. Azure is best for teams already invested in the Microsoft ecosystem who want the conversational language understanding features. For cost control strategies, see our Free API Cost and Quota Control in Practice page.
9. FAQ
Q: What is the difference between slot filling and NER? A: NER only discovers entities. Slot filling additionally determines which field each entity belongs to and whether the field value is valid.
Q: Can free APIs handle multi-turn conversations? A: The free tiers of Hugging Face and Google Cloud are stateless single-call APIs. Multi-turn conversation context must be maintained by your own application logic.
Q: How accurate is Chinese slot filling? A: Hugging Face hosts numerous Chinese fine-tuned models that achieve near-commercial accuracy on Chinese entity recognition benchmarks. Models like uer/roberta-base-finetuned-chinaner-chinese and hfl/chinese-macbert-base typically score above 90 F1 on MSRA and OntoNotes datasets. For domain-specific scenarios such as healthcare or finance, further fine-tuning on domain corpora typically requires only a few thousand annotated examples to achieve significant improvements.
Q: Where should I start? Browse our Free API Cost and Quota Control in Practice to select a channel, and use our Free Text Summarization API Complete Tutorial registration link to experience the complete solution. We recommend starting with Hugging Face: register for an account, obtain an Access Token, and run through the first slot filling example in under five minutes. Once validated, migrate to other channels based on actual requirements. For teams with strict latency requirements, consider deploying a local spaCy model with custom entity rulers. For teams with complex schema validation, consider building a rule layer on top of any of the free NER APIs. The key insight is that slot filling is not a single API call — it is a pipeline that combines extraction, mapping, and validation. Start simple, measure honestly, and iterate based on real user queries. The first version of your slot filling system should use a single NER model with a basic rule layer. Measure its accuracy on a representative sample of real user inputs. Identify the most common failure modes — typically missing slots, wrong slot assignments, or values that fail validation — and address them one at a time. As your accuracy improves, you can add more sophisticated ML models for harder cases. The goal is not to achieve perfect accuracy from day one, but to build a system that improves continuously as you collect more data and refine your rules and models. For teams with no ML expertise, the rule-based approach alone can achieve 85–90% accuracy on well-defined dialog flows, which is sufficient for many production applications. We recommend starting with Hugging Face: register for an account, obtain an Access Token, and run through the first slot filling example in under five minutes. Once validated, migrate to other channels based on actual requirements.