← Back to articles
Tutorials

Free Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)

Free Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)

The one-sentence verdict: if you want to turn unstructured text into structured entity labels (people, places, organizations, money, dates) without paying for token metering or GPU time, the free tiers of Azure AI Language, Amazon Comprehend, Google Cloud Natural Language, plus open-source spaCy, are more than enough for most information extraction, social listening, and knowledge graph use cases. This guide verifies the free quotas and setup steps of all four channels as of 2026-10-04, with commands you can reproduce directly.

1. What NER Is and Why It Matters

Named Entity Recognition (NER) is one of the most immediately useful abilities in natural language processing. Given a piece of text, it automatically identifies proper nouns and tags them with categories — people (PER), locations (LOC), organizations (ORG), money (MONEY), dates (DATE), times (TIME), percentages (PERCENT), and so on.

For example, feeding "Apple CEO Tim Cook announced in San Francisco yesterday that third-quarter revenue hit $90 billion" into NER produces:

  • "Apple" → ORG
  • "Tim Cook" → PER
  • "San Francisco" → LOC
  • "yesterday" → DATE
  • "$90 billion" → MONEY

This is not simple keyword matching. NER models understand context: they know whether "Apple" means the fruit or the company, and whether "Washington" is a person or a city. That contextual understanding is what makes NER dramatically more powerful than hand-written regular expressions, and why it feeds so many downstream tasks. You can find this and other free capabilities collected in one place in our Free API hub, without signing up for each service separately.

2. How Modern NER Works

Early NER relied on rules and dictionaries: prepare a gazetteer of names, places, and organizations, then do string matching against the text. It is fast but brittle — anything not in the dictionary gets missed, and ambiguous words like "Apple" cannot be disambiguated.

Modern NER is dominated by deep learning, roughly in two flavors:

  1. Statistical sequence labeling: the sentence is tokenized, and a model predicts for each token whether it is the beginning of an entity (B-), inside an entity (I-), or outside (O). These models are lightweight and fast, ideal for real-time calls.
  2. Large language model extraction: you prompt an LLM to output structured entity JSON directly. It is more capable and understands complex context, but costs more on long documents and needs multiple calls.

Under the hood, the sequence-labeling approach is usually a bidirectional transformer (BERT and its descendants) with a classification head on top. During training, the model learns to attend to the surrounding words so it can resolve ambiguity — "Apple" followed by "announced" is a company, while "Apple" followed by "is sweet" is a fruit. This is also why multilingual models need per-language training data: the surrounding cues that disambiguate entities differ from one language to another.

For most free-tier workloads, a vendor's pretrained NER (essentially a mature sequence labeling model) is plenty, and you never have to train a model yourself. For the free quota of the LLM-based route, see the free model list in our Free LLM API Rankings.

3. Four Realistic Use Cases

The free capabilities needed for the following scenarios all have reproducible tutorials in our Free API hub, so you don't need to reinvent anything.

  1. Social listening: automatically extract which company, product, or person is mentioned across a stream of news, then aggregate mention counts for brand tracking. If you ingest 1,000 articles daily, NER pulls out the hottest brand and person names.
  2. Financial intelligence: extract amounts, dates, and counterparties from earnings reports and announcements to support risk and compliance review. A single report carries dozens of figures — NER structures them in one pass.
  3. Knowledge graphs: structure entities and relationships before feeding them into a graph database for reasoning and Q&A. Entity recognition is step one of building the graph; without entities there are no nodes.
  4. Support ticket triage: identify order numbers, locations, and product models mentioned in tickets to route them automatically and cut manual triage time.

Each of these starts with the same primitive — extract entities, then post-process. The rest is plumbing, which is why getting the extraction layer right (and cheap) matters more than any downstream modeling decision you will make later.

4. Four Free NER Channels Compared (Verified 2026-10-04)

Here is a side-by-side of the mainstream free NER channels in 2026:

Channel Free quota Credit card required Entity types Key strength
Azure AI Language 5,000 transactions/month No (F0 free tier) 4 core types, multilingual Microsoft official, high pretrained quality
Amazon Comprehend 50K units (5M characters)/month/API Yes (AWS account) 12 types, incl. PII Widest coverage, dedicated PII detection
Google Cloud Natural Language 5,000 units/month (permanent) No (GCP project) Multilingual Permanent free + $300 new-account credit
spaCy (open source) Fully free, no limit No Dozens of language models Runs locally, data stays on-prem

A note on units: Amazon Comprehend bills in "units" where one unit equals 100 characters, and a "free tier" covering 50K units per API per month amounts to 5 million characters — generous for a month of personal or small-team use. Azure and Google bill transactions/units per request, and counting carefully matters because a request that is 90% empty space costs the same as a full one.

Quota figures verified against vendor documentation as of 2026-10-04; always check the console for changes. Two key differences: first, the three cloud services bill by request count or character units with a free quota ceiling, while spaCy is fully free to run locally but requires self-hosting; second, only Amazon Comprehend provides dedicated PII (sensitive information) detection.

5. Five-Minute Setup (Azure AI Language, tested)

Azure AI Language's NER is a prebuilt capability — no model training needed. Once you have a key, you call it directly:

  1. In the Azure Portal, create a "Language service" resource and pick the F0 (free) pricing tier.
  2. Copy the resource Endpoint and Key (the key is shown only once after creation).
  3. Send text with the Python snippet below to get entity annotations back.

One subtlety worth underlining: the F0 free tier and the paid S tier share the same API surface, but the free tier has stricter throttling and a lower monthly ceiling. If you accidentally leave an S-tier resource running, you will pay per transaction even for empty traffic, so double-check the pricing tier at creation time — this is a common "surprise bill" trap for newcomers.

Free-tier test notes: the free tier caps a single request at 5,120 characters of input and the monthly limit is 5,000 transactions; exceeding it returns 429. Chunking long text into sub-5,120-character pieces and stuffing each one full is far more economical than many small requests.

6. Python Call (the only code sample)

Using the official Azure AI Language SDK, entity recognition is just a few lines:

# pip install azure-ai-textanalytics
from azure.ai.textanalytics import TextAnalyticsClient
from azure.core.credentials import AzureKeyCredential

client = TextAnalyticsClient(
    endpoint="https://<your-resource>.cognitiveservices.azure.com/",
    credential=AzureKeyCredential("<your-key>"),
)

docs = ["Apple CEO Tim Cook announced in San Francisco that Q3 revenue hit $90 billion."]
result = client.recognize_entities(docs)[0]

for entity in result.entities:
    print(f"{entity.text} -> {entity.category} (confidence {entity.confidence_score:.2f})")

Key points: every returned entity carries text (the original), category, and confidence_score (higher is better). The recognize_entities call accepts a batch of documents, so loop over chunks and call once per batch rather than once per sentence to stay well inside the transaction quota. In production, manually review results below a 0.5 confidence threshold to avoid mislabeling "Apple" as a company when it is a fruit.

7. Entity Type Comparison Across the Three Clouds

Each channel's predefined entity types differ slightly, so scan this table before choosing:

Entity type Azure AI Language Amazon Comprehend Google Cloud NL
Person Person PERSON PERSON
Organization Organization ORGANIZATION ORGANIZATION
Location Location LOCATION LOCATION
Money/Quantity Quantity QUANTITY MONEY
Date/Time DateTime DATE DATE
Other Product, event, etc. 12 types incl. PII, event Event, work, etc.

Amazon Comprehend's twelve types are the widest of the three and are the only ones that natively cover PII (personal identifiable information), which wraps identifiers like national ID numbers, bank cards, email addresses, phone numbers, and driver's licenses. This is a compliance feature rather than a general NLU feature, and it is the most easily overlooked difference when you choose a channel for a medical, financial, or HR document pipeline.

When you need sensitive ID detection, Amazon Comprehend's Detect PII is purpose-built — Azure's and Google's general NER do not cover this.

8. Running spaCy Locally (No API Key, No Quota)

If you want a fully free, offline option, spaCy is the workhorse. It ships pretrained pipelines for English and more than a dozen other languages, and the EntityRecognizer runs entirely on your own machine.

Setup is two steps: install the library, then download a pipeline. The en_core_web_sm model is small and fast; en_core_web_trf is a transformer-based model with higher accuracy but a bigger download and slower runtime. For a production Chinese pipeline you would use zh_core_web_sm or zh_core_web_trf.

A minimal run looks like this (conceptually):

  • Load the pipeline with spacy.load("en_core_web_sm").
  • Pass your document through the pipeline.
  • Iterate over doc.ents and read each entity's .text and .label_.

Because spaCy is local, three things matter more than with cloud APIs: model size vs accuracy (the small models are fast but miss rare entities), memory footprint (the transformer models want several gigabytes of RAM), and pipeline speed (disable the components you don't need, like the parser, to accelerate batch extraction). There is no monthly quota and no per-call cost — the trade-off is that you maintain the environment yourself and every document stays on your infrastructure, which is exactly what compliance-sensitive teams want.

For teams already inside a Python data stack, spaCy is often the zero-friction default: no account, no billing, no vendor lock-in, and the same code runs in a notebook, a cron job, or a containerized microservice.

9. Calling Amazon Comprehend and Google Cloud NL

The other two cloud channels follow the same pattern as Azure — create a resource, grab credentials, call an SDK or REST endpoint — with a few provider-specific quirks.

Amazon Comprehend exposes DetectEntities for general entity recognition and DetectPiiEntities for sensitive data. It requires an AWS account and credit card even for the free tier, and billing is per 100-character unit. The free tier is 50K units per API per month, which is generous but shared across the eligible APIs, so a heavy sentiment workload can eat into your entity budget. Use the BatchDetectEntities and async StartEntitiesDetectionJob endpoints when processing more than a handful of documents in one go — the sync endpoint is meant for single requests and will feel slow on a long queue. The response returns each entity's score, type (PERSON, ORGANIZATION, LOCATION, COMMERCIAL_ITEM, DATE, and others), and begin/end offsets you can slice directly from the input.

Google Cloud Natural Language exposes analyzeEntities through the language client. Its free tier is 5,000 units per month permanently (each analyzeEntities call is one unit for the standard API), plus a one-time $300 new-customer credit that applies to the paid feature set. Google's entity response distinguishes entity mentions from entity names, gives you a salience score (how central the entity is to the document), and can link well-known entities to a Knowledge Graph entry via metadata and mid fields. That salience score is uniquely useful for summarization and social-listening use cases — it tells you which entity the writer actually cares about, not just which entity appears most often.

The practical decision usually comes down to: Chinese quality → Azure, widest categories + PII → Amazon, permanent free + salience → Google, offline/privacy → spaCy. Most small teams run one primary cloud channel plus spaCy as a local fallback, which covers nearly every scenario at zero marginal cost.

10. Seven Pitfalls (read these before anything else)

  1. Free tiers have character caps: Azure F0 allows 5,120 characters per request, Amazon has a 300-character minimum per request — chunk long text first, then stitch results together.
  2. Quota is per request, not per entity: Azure counts one transaction per request, so one full 5,120-character request is far better than ten small ones.
  3. Set a confidence threshold: general NER gets rare entities wrong (niche brands, dialect place names), so keep a human fallback below the threshold.
  4. Match language to model: use a Chinese model for Chinese text and an English model for English; mixing them noticeably degrades accuracy.
  5. Retry with backoff on throttling: cloud services return 429 at peak hours, so add exponential backoff (three retries at 1s/2s/4s).
  6. Keep sensitive data offline: for privacy or compliance text, prefer running spaCy locally so data never leaves your network.
  7. Free quotas change: vendor free-tier policies shift frequently, so confirm against the official pricing page before production.

Two more worth internalizing: offsets are not always UTF-8 — some APIs return character offsets in a specific encoding (UTF-16 code units for Azure/AWS), so if you slice the original string with the wrong offset convention you will clip entities mid-character on multibyte text. And async vs sync: Amazon Comprehend offers batch and async endpoints for large jobs; using the sync endpoint for a thousand documents will be slow and messy, while the async endpoint processes them in one place and returns results to an S3 bucket.

Test data verified through 2026-10-04. Free policies change often; always confirm against official vendor docs before production use.

11. Measuring Quality (Precision, Recall, F1)

If you evaluate vendors or your own extraction pipeline, use the three classic metrics:

  • Precision: of the entities you returned, how many are correct. High precision means few false positives.
  • Recall: of the true entities in the text, how many you found. High recall means few false negatives.
  • F1: the harmonic mean of the two, the single number most people optimize.

For general English news text, pretrained NER models routinely land in the 0.85–0.95 F1 range; on messy user-generated content or rare domain jargon, expect a drop that proper thresholds and post-processing can partially recover. When you compare vendors, always measure on your own domain text, not a generic benchmark — a model that wins on CoNLL may lose on your medical records.

12. FAQ

  • Is the free quota enough? With Azure's 5,000 requests/month as an example, processing 150 articles a day is roughly 4,500 requests a month — more than enough; most desktop projects never exhaust it.
  • Which is best for Chinese? Azure AI Language's Chinese pretrained model is stable and is the top pick for Chinese scenarios; choose Amazon for the widest categories, Google for permanent free tier.
  • How do I keep entity offsets straight across multibyte text? Read the API's documented offset convention (Azure and AWS use UTF-16 code units for some endpoints) before slicing the original string, otherwise you will clip entities mid-character. Then, separate from that, here's the upload question — what if I don't want to upload data? Run spaCy locally — after downloading the model it is fully offline and data never leaves your machine, ideal for compliance-sensitive scenarios. Other free text capabilities are listed in the Free API hub.

13. Unify NER with the Rest of Your Text Stack

After entity recognition you usually need follow-up processing — keyword extraction for topics, semantic similarity for dedup, RAG for knowledge base search. Every one of these has a free recipe in our Free API hub:

Think of NER as the extraction layer and these other tools as the processing layer: NER hands you the "who, where, when, how much," then keyword extraction summarizes the "what," semantic similarity collapses duplicates, and a vector database plus reranker turn the whole thing into a searchable knowledge graph. Chaining them with free tiers keeps the entire pipeline at zero marginal cost until you scale past the quotas.

14. A Worked Cost-and-Scale Example

To make the quotas concrete, imagine a small content team that wants to extract entities from 3,000 support tickets a month, each averaging 800 characters. Here is how the four channels compare against that workload:

  • Azure AI Language (F0): 3,000 tickets ≈ 2.4 million characters; at a 5,120-character request cap you need roughly 500–600 batched requests, comfortably inside the 5,000-transaction monthly free tier. Zero cost.
  • Amazon Comprehend: 2.4 million characters = 24,000 units, under the 50K-unit monthly free tier. Zero cost, but you paid the upfront friction of an AWS account with a credit card.
  • Google Cloud Natural Language: 500–600 analyzeEntities calls, under the 5,000-unit permanent free tier. Zero cost, with the $300 credit still untouched for heavier experiments.
  • spaCy: no quota at all, but you run the compute yourself. On a modest machine the en_core_web_sm model handles a few thousand short documents in minutes; the transformer model is more accurate but visibly slower and RAM-hungry.

The pattern is clear: for small-to-medium volumes, every channel is effectively free, and the choice is driven by category coverage, language quality, and whether data may leave your network rather than price. The moment you cross into hundreds of thousands of documents a month, the cloud free tiers run out and you start pricing per character/unit — at which point moving to self-hosted spaCy (or a batched async pipeline) becomes a cost decision rather than a convenience one.

15. Building a Minimal Extraction Pipeline

A practical pipeline has four stages, and each stage is where beginners silently misconfigure things:

  1. Ingest and normalize: pull text from your source (CSV, API, scraper) and strip markup, HTML, and control characters. NER models are trained on clean text; feeding them raw HTML with tags and encodings measurably hurts accuracy.
  2. Chunk and batch: split long documents at sentence boundaries (never mid-word) to respect each API's character cap, then group chunks into the largest legal batch so you spend the fewest transactions.
  3. Extract and threshold: call the entity endpoint, then filter by confidence. Low-confidence entities go to a human review queue; everything above threshold flows straight through.
  4. Deduplicate and normalize: NER returns the same entity with different surface forms ("Apple", "Apple Inc.", "AAPL"). Normalize with a canonical name and merge duplicates so your downstream counts aren't inflated.

The fourth stage is the one teams underestimate: without normalization, a social-listening report will list "Apple", "Apple Inc", and "Tim Apple" as three separate entities, and your mention counts will be garbage. A simple dictionary or an embedding-based clustering step fixes this, and it pairs naturally with the other free text tools covered on this site.

16. When NER Alone Isn't Enough

NER tells you the "who," "where," and "how much," but not the "what they did" or "how they feel about it." When you need those layers, reach for the adjacent capabilities:

  • Sentiment analysis to score whether a mention is positive or negative — essential for brand health tracking beyond raw counts.
  • Keyword extraction and topic modeling to summarize what each cluster of mentions is about.
  • Semantic similarity to deduplicate near-identical documents before counting, so one viral story isn't counted fifty times.
  • Relation extraction (a step up from NER) to link two entities together, such as "works at" or "acquired."

Each of these has a free-tier recipe in the apishare.cc catalog, and chaining them turns a bare entity list into an actual intelligence pipeline. The mental model to keep is simple: NER gives you the entities, the neighboring tools give you the relationships and the meaning, and together they cover the vast majority of what smaller teams actually need.

17. A Quick Decision Checklist

Before you build, answer five questions and the channel choice follows almost mechanically:

  • Does the data contain sensitive identifiers? If yes, start with Amazon Comprehend's Detect PII — no other free general NER covers it natively.
  • Is your primary text Chinese or multilingual? Azure AI Language's pretrained Chinese model is the most consistently reliable free option.
  • Must data never leave your network? Then spaCy local is non-negotiable, full stop.
  • Do you need salience or Knowledge Graph linking? Google Cloud Natural Language is the only one of the three that returns a salience score and entity mid out of the box.
  • Are you already on AWS or GCP? If you already have an account and billing set up, the corresponding native service is the lowest-friction path — the free tier starts immediately with zero new signup.

Answer those five, and you will spend minutes rather than days choosing. The good news is that all of these routes cost nothing at small scale, so the risk of picking "wrong" is low — you can prototype against one channel today and switch with a thin abstraction layer whenever your scale or requirements change.

Claim Your Free Quota and Keep Reading

Want to call all the free models above with one unified key and no per-provider signups? Apishare.cc provides a single API key that reaches 100+ models, with free models costing zero. New users can grab their free quota on the registration page, and find more tutorials and rankings in the Free API hub.

More in this category

Free Text Summarization API Complete Tutorial: Let LLMs Compress 1M-Word Documents into 100 WordsFree Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)Free Time Series Forecasting API Complete Tutorial: Zero-Cost “Crystal Ball” for Sales/Inventory/Energy Prices (Verified 2026-10-03)Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02)Free Voice Cloning API Complete Tutorial: Clone Your Signature Voice from a Reference Clip

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.