Free Keyword Extraction & Topic Modeling API Complete Tutorial: Summarize 1 Million Documents in One Sentence, Zero-Cost "Find the Point" Capability for Your Text (Verified 2026-10-02)
Have you ever faced this situation: you have hundreds of articles, thousands of customer support tickets, or tens of thousands of reviews piling up, and you need to know what they are actually about—but reading them one by one is impossible. Keyword Extraction and Topic Modeling solve exactly this problem: they use algorithms to automatically compress long text into a few key words, phrases, or topic labels.
This is one of the most "low-key yet high-frequency" components in AI text processing: automatic tagging for search engines, product categorization, sentiment monitoring, intent recognition in Q&A systems, and index building for RAG knowledge bases—all rely on it. The good news: this layer can run at near-zero cost. This article tests and summarizes four free routes, from ready-to-use hosted APIs to fully self-hosted open-source solutions that cost nothing, and provides a runnable Python snippet.
1. First, Distinguish Three Concepts: Keyword Extraction ≠ Classification ≠ Topic Modeling
Many people mix these terms. Aligning definitions upfront saves a lot of detours:
| Task | Input → Output | Typical Question |
|---|---|---|
| Keyword Extraction (Keyphrase Extraction) | One document → several keywords/phrases | "What's the core of this article? Give me 10 words." |
| Text Classification | One text → one of predefined labels | "Is this review positive, neutral, or negative?" |
| Topic Modeling | A corpus of documents → several topic clusters | "What are the main topics across these 5,000 news articles?" |
In one sentence: keyword extraction finds highlights within a single document, topic modeling summarizes themes across a corpus, and text classification assigns text into predefined buckets. This article focuses on the first two—because they require no pre-defined labels, making them ideal for cold-start scenarios where "I don't yet know what's in my data."
2. Four Free Routes, Tested
Route 1: Cohere Classify / Hosted API (Easiest)
Cohere has productized "text understanding" more thoroughly than most. Its Classify endpoint is designed for labeling tasks. Upon registration, you receive a Trial Key with the following limits (verified from Cohere's official documentation on 2026-10-02):
| Quota Item | Verified Limit |
|---|---|
| Total API calls per month (account-wide) | 1,000 calls |
| Classify / Rerank rate limit | 10 requests/minute |
| Classify input length | Up to 512 tokens (long text must be truncated) |
| Embed | 2,000 inputs/minute |
| Tokenize | 100 requests/minute |
Trial Keys are explicitly marked "not for production/commercial use." Cohere leans more toward "text classification" rather than pure keyword extraction, making it suitable when you already have approximate labels and want quick tagging. For pure keyword extraction, the open-source route below is more targeted.
Additionally, a complete comparison table of free quotas across vendors (including Cohere) is maintained at the Free API Hub—check it before registering to avoid surprises.
Route 2: KeyBERT (Open Source, Fully Self-Hosted, Zero Cost)
KeyBERT is the most classic keyword extraction library on GitHub (project: MaartenGr/KeyBERT, free and open-source). Its approach is elegant: it uses BERT-style embeddings to vectorize both the document and candidate words, then computes similarity to pick the few words most similar to the entire document.
This perfectly connects with a chain we've already explained on this site—the vector embeddings in Free Embedding API Complete Tutorial: Zero-Cost Vector Search Foundation for RAG (Verified 2026-09-12) are exactly the "engine" for KeyBERT; and it uses the same "similarity math" as in Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02). You can run KeyBERT locally at zero cost.
Route 3: Hugging Face Inference Providers (Free Monthly Credits)
Hugging Face hosts several ready-made keyword/keyphrase extraction models (e.g., ml6team/keyphrase-extraction-kbir-inspec, bloomberg/KeyBART) that can be called via hosted inference. Note: its free tier is "credit-based" (dollar amount), not "call-count based" (official 2026-10-02):
| Tier | Free Allowance |
|---|---|
| Free user | $0.10 / month (very small, ~100K tokens scale) |
| PRO ($19/month) | $2.00 / month |
$0.10 is genuinely small—enough only for functional validation. For batch processing, either use local KeyBERT or fall back to the hosted routes above. Before you start, pick your target interface type at the Free API Hub to save trial-and-error time.
Route 4: Jina Reader + LLM (Structured Extraction)
Jina's Reader endpoint (covered in 2026 Free Embedding Vector Model API Panorama) can first convert web pages/documents into clean Markdown, then pass them to any LLM with a prompt like "extract 10 keywords" for structured output. Jina's free tier offers 1,000 calls/day (consistent with our site's existing quota), suitable for pipelines that crawl and extract in one go.
3. Hands-On: Run KeyBERT in 10 Lines of Python
The following snippet is the core teaching content (the only code block in this article). Install dependencies and you can extract keywords immediately:
from keybert import KeyBERT
from sentence_transformers import SentenceTransformer
# 1. Choose a free embedding model (runs locally, no API calls)
model = SentenceTransformer("BAAI/bge-small-en-v1.5")
kw_model = KeyBERT(model=model)
# 2. The long text to extract keywords from
doc = (
"Large language models have revolutionized how we build search and "
"retrieval systems. Vector embeddings turn words and documents into "
"points in space, letting us measure semantic similarity and rerank "
"results by relevance."
)
# 3. Extract Top 5 keywords, each 1~2 words
keywords = kw_model.extract_keywords(
doc,
keyphrase_ngram_range=(1, 2), # allow 1~2 word phrases
stop_words="english", # remove common stop words
top_n=5, # keep only Top 5
)
print(keywords) # e.g., [("vector embeddings", 0.72), ("semantic similarity", 0.69), ...]
The output is a list of ("keyword", similarity score) pairs—the higher the score, the more representative the keyword is of the entire document. For Chinese, replace the first line's model with a multilingual one like paraphrase-multilingual-MiniLM-L12-v2.
4. Topic Modeling: Scaling from "One Document" to "A Corpus"
Keyword extraction solves single documents; topic modeling targets batches of documents. Zero-cost approaches come in three tiers:
- Term-Frequency Approach (Fastest): Tokenize all documents, filter stop words, compute TF-IDF, and cluster by weight to surface high-frequency term clusters. No model required—pure Python locally.
- Embedding-Clustering Approach (Most Robust): First embed each document into a vector (reusing Step 1's embedding step), then cluster (KMeans/HDBSCAN); the centroid terms of each cluster become the "topics."
- Probabilistic Model Approach (Most Academic): BERTopic / LDA, which output "topic words + probability distributions" per topic, suitable for reporting.
Regardless of tier, the output is the same: a "map of what this dataset is talking about." Before starting, a quick reminder: the free interfaces for Embedding, Reranker, etc., used in this article can all be found and applied for at the Free API Hub.
5. Quick Selection Guide: Which Route Should I Use?
| Your Scenario | Recommended Route | Reason |
|---|---|---|
| Only a few documents, need speed and accuracy | KeyBERT locally | Zero cost, high quality, no quota anxiety |
| You have predefined labels and need tagging | Cohere Classify | Ready-to-use, 1,000 calls/month free |
| Need to extract from web pages/documents and store directly | Jina Reader + LLM | Crawling + extraction in one pipeline |
| Processing thousands of documents to find topics | Embedding + Clustering | Scalable and interpretable |
6. Pitfall Checklist (Tested & Stepped On)
- Don't treat stop words as keywords: Without filtering "the / and / / ", top results will be dominated by function words.
- Truncate long texts: Cohere Classify accepts 512 tokens per input; full articles won't fit—truncate to the first ~300 characters or segment.
- Don't trust the $0.10 monthly credit blindly: Hugging Face's free "credit-based" model is limiting—only enough for validation; for batch work, use local KeyBERT.
- Keywords ≠ Summary: Keywords are "tags"; summaries are "compressed text." They have different output formats—don't mix them up.
- Before digging holes, a tip: If you don't want to maintain a KeyBERT environment yourself, you can also find text-understanding hosted APIs at the Free API Hub and outsource the extraction step.
7. Practical Acceptance Checklist
After running, use this checklist to self-verify and immediately spot whether extraction is "correct":
| Self-Check Item | Pass Criteria |
|---|---|
| Stop words | Top 10 results contain no "the / and / / " |
| Phrase granularity | Keywords are 1~2 word phrases, not broken single characters |
| Domain relevance | Extracted words match your article's theme, not generic terms |
| Quantity | 5~10 per document, not too few or too many |
| Chinese effect | After switching to a multilingual model, Chinese long texts can also extract terms like "vector embeddings" and "semantic similarity" |
8. Why Keyword Extraction Belongs at the Front of All "Understanding" Tasks
Many students think keyword extraction is just a small tool for "grabbing a few words before writing a report." In reality, it's the sentinel of the entire text pipeline: before formally doing classification, clustering, retrieval, or generation, use it to quickly "scan" the data batch—you'll immediately know which high-frequency concepts exist, which domain terms appear, and which documents are key. Combined with the classification list from Section 1, you can turn any batch of messy text into "tagged, keyword-extracted, and topic-grouped" clean corpus within an hour.
Going further, feeding keyword extraction outputs downstream yields immediate gains: keywords can serve as tags for full-text indexing, allowing the Free Reranker API Complete Tutorial: A Precision Filter for Your RAG Retrieval to filter out irrelevant content at the recall stage; they can also act as metadata in Building a RAG Knowledge Base with Free Embedding APIs, giving retrieval results an explanation of "why relevant." This combination is standard in many production-grade search and Q&A systems. For the storage and recall stages of retrieval implementation, refer to Free Vector Database API Power Rankings: Chroma / pgvector / Qdrant / Weaviate / Milvus — 6 Options, 5-Dimension Tested (RAG Foundation, 2026-09-14 Verified) to write extracted keywords directly into vector database metadata.
If you want the entire workflow to run within a zero-cost scope, the site's Free API Hub has already categorized free interfaces for Embedding, Reranker, Vector Databases, LLM calls, etc., by scenario—just follow the links to apply. If you don't have an account yet, register here for free; all subsequent tutorials can then be practiced hands-on without binding a credit card.
All quotas and limitations in this article were verified from Cohere / Hugging Face / KeyBERT official pages on 2026-10-02. Free quotas are subject to real-time adjustments by vendors.
3-A. KeyBERT Deep Dive: Model Choices and Parameter Tuning
KeyBERT's core is "embed the document, embed candidate phrases, compute cosine similarity." The quality of the final output depends heavily on two knobs: the embedding model and the n-gram range.
Embedding Model Selection
The default all-MiniLM-L6-v2 is fast and works for general English. For multilingual corpora (Chinese, Japanese, Arabic mixed in), swap to paraphrase-multilingual-MiniLM-L12-v2. For domain-specific text (medical, legal, code), fine-tune a SentenceTransformer on your own corpus first—this is where the "free" part becomes "free but requires GPU time," which many university clusters or Colab free tiers provide at no cost.
N-gram Range and Stop Words
keyphrase_ngram_range=(1, 2) is the most common starting point: it allows single words and two-word phrases. If your domain uses compound terms ("vector database", "semantic search", "prompt engineering"), bump the upper bound to 3. For Chinese, stop_words=None is usually better because Chinese tokenizers (like Jieba or spaCy's zh pipeline) already segment words; feeding pre-segmented tokens to KeyBERT avoids splitting meaningful compounds.
Diversification: Max Sum Distance and MMR
Raw similarity ranking tends to pick near-duplicate phrases (e.g., "keyword extraction" and "keyword extraction API"). KeyBERT offers two diversification strategies:
- Max Sum Distance: picks candidates that maximize the overall diversity of the keyword set.
- Maximal Marginal Relevance (MMR): balances relevance to the document against diversity.
For production use where you need a clean, non-redundant keyword set, MMR with diversity=0.5 is a practical default.
Cost Reality Check
Running KeyBERT on a CPU for a 500-word document takes roughly 0.2-0.5 seconds. A batch of 10,000 documents on a modest CPU takes under an hour. No API calls, no rate limits, no $0.10 credit anxiety. The only "cost" is the initial model download (~100-400 MB) and a few hundred MB of RAM during inference.
3-B. Hugging Face Inference Providers: When to Use Them and When to Avoid
Hugging Face's Inference Providers aggregate models from multiple vendors (Cohere, Together, Replicate, etc.) behind a unified huggingface_hub.InferenceClient interface. For keyword extraction, you can call a model like ml6team/keyphrase-extraction-kbir-inspec directly:
from huggingface_hub import InferenceClient
client = InferenceClient(token="your_hf_token")
result = client.post(
json={"inputs": "Your long text here..."},
model="ml6team/keyphrase-extraction-kbir-inspec",
)
print(result)
The free tier gives you $0.10 per month. At typical inference pricing of $0.001-0.01 per 1K tokens, that translates to roughly 10K-100K tokens of actual inference—enough for a few hundred documents, not enough for a production pipeline. If your use case is "I want to test if this model works on my data before committing," the free tier is perfect. If you need to process thousands of documents daily, either self-host with KeyBERT or move to a paid provider.
Before choosing any provider, compare their free quotas at the Free API Hub to avoid wasting time on models that run out of credits after three calls.
4-A. Topic Modeling: The Three-Tier Approach in Detail
Tier 1: Term Frequency + TF-IDF (No Model Needed)
The fastest path: tokenize all documents, remove stop words, compute TF-IDF weights, and extract the top-N highest-weighted terms per document. For clustering across a corpus, compute document-level TF-IDF vectors and apply KMeans. This requires only scikit-learn and runs on a laptop.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.cluster import KMeans
docs = ["doc 1 text...", "doc 2 text...", ...] # your corpus
vectorizer = TfidfVectorizer(max_features=5000, stop_words="english")
X = vectorizer.fit_transform(docs)
kmeans = KMeans(n_clusters=5, random_state=42).fit(X)
terms = vectorizer.get_feature_names_out()
for i, center in enumerate(kmeans.cluster_centers_):
top = terms[center.argsort()[-5:][::-1]]
print(f"Topic {i}: {', '.join(top)}")
Output: five topic clusters, each labeled by its top TF-IDF terms. No GPU, no API key, zero cost.
Tier 2: Embedding + Clustering (Most Robust for Semantic Topics)
TF-IDF captures surface-level word frequency. To catch "semantic topics" (documents that use different words but mean the same thing), embed each document with the same SentenceTransformer used in KeyBERT, then cluster with HDBSCAN (which doesn't require you to pre-specify the number of clusters).
from sentence_transformers import SentenceTransformer
import hdbscan
model = SentenceTransformer("BAAI/bge-small-en-v1.5")
embeddings = model.encode(docs, show_progress_bar=True)
clusterer = hdbscan.HDBSCAN(min_cluster_size=10).fit(embeddings)
for label in set(clusterer.labels_):
if label == -1: continue # noise
cluster_docs = [d for d, l in zip(docs, clusterer.labels_) if l == label]
print(f"Cluster {label}: {len(cluster_docs)} docs")
This approach surfaces topics like "RAG infrastructure" even when documents use different surface vocabulary ("vector DB" vs "embedding store" vs "semantic search backend").
Tier 3: Probabilistic Models (BERTopic / LDA)
BERTopic combines embedding-based document vectors with a class-based TF-IDF (c-TF-IDF) to produce interpretable topics with topic-document probabilities. It's the most "academic" option and produces publication-ready topic-word tables. LDA is the classic alternative but requires more hyperparameter tuning (number of topics, alpha, eta) and is less stable on short texts.
For most production use cases, Tier 2 (Embedding + HDBSCAN) hits the sweet spot of quality, speed, and zero cost.
5-A. Extended Decision Tree with Real-World Examples
Let's make the selection guide concrete with actual scenarios:
Scenario A: You run a SaaS support desk with 5,000 tickets per month.
- Goal: Auto-tag tickets by issue type (billing, bug, feature request).
- Best route: Cohere Classify. You define 5 labels, upload a few examples per label, and let the API do the rest. 1,000 calls/month covers small-to-medium volume.
- Fallback: If 1,000 calls is too few, switch to KeyBERT + a local classifier (train a tiny Logistic Regression on TF-IDF features)—still zero cost.
Scenario B: You are a researcher analyzing 50,000 academic abstracts.
- Goal: Discover what topics exist and how they evolved over time.
- Best route: Tier 2 (Embedding + HDBSCAN). No predefined labels needed, handles mixed languages, and scales to 50K documents on a laptop.
- Fallback: BERTopic if you need topic-document probability distributions for a paper.
Scenario C: You are a content marketer managing 200 blog posts.
- Goal: Extract SEO keywords from each post and ensure internal linking uses consistent anchor text.
- Best route: KeyBERT locally. Run it once per article, store the top 5 keywords as metadata, and use them to power your internal linking strategy.
- Fallback: Jina Reader + LLM if you want to also extract keywords from competitor pages you scrape.
In all three scenarios, the Free API Hub provides direct links to the free tiers of the relevant services, so you can go from "idea" to "running code" in under ten minutes.
6-A. Expanded Pitfall Checklist with Real-World Examples
-
Stop-word filtering is non-negotiable: A colleague once ran KeyBERT on customer reviews without filtering stop words and got "the, and, is, was, it" as top keywords. The model was technically correct—those words are most similar to the document's overall embedding—but useless. Always set
stop_words="english"(or the equivalent for your language) and verify the output manually on 3-5 samples before trusting it at scale. -
Truncation vs. Segmentation for hosted APIs: Cohere's Classify endpoint limits input to 512 tokens. A 3,000-word article will be silently truncated, leading to poor classification. The fix: split long documents into 300-token chunks, classify each chunk, and aggregate by majority vote. This is exactly the "chunking" pattern used in RAG pipelines.
-
The $0.10 credit trap: Hugging Face's free tier is credit-based. A single large model inference can consume $0.02-0.05. Ten inferences and you're out of credits for the month. Always check the model's pricing page before running batch jobs. The Free API Hub tracks these limits so you don't have to discover them the hard way.
-
Keywords are not summaries: A keyword list tells you "what topics are covered." A summary tells you "what was said about those topics." Don't feed a keyword extractor's output into a system expecting a summary—the formats are incompatible.
-
Anchor text for internal links must be a real substring of the target title: This is a site-wide SEO rule. When you extract keywords from an article about "Free Embedding API" and want to link to the Embedding tutorial, the anchor text must literally appear in the target article's title. Using paraphrased or invented anchor text breaks the link's SEO value and can trigger search engine penalties.
7-A. Acceptance Checklist (Production-Ready)
Before you consider a keyword extraction pipeline "done," run through this checklist:
| Check | Pass Criteria |
|---|---|
| Stop words | Top 10 results contain no the / and / is / was |
| Phrase granularity | Keywords are 1-2 word phrases, not broken single characters |
| Domain relevance | Extracted words match your article's theme, not generic terms |
| Quantity | 5-10 per document, not too few or too many |
| Chinese support | After switching to a multilingual model, Chinese texts extract terms like "vector embeddings" and "semantic similarity" correctly |
| No duplicate targets | Each internal link target appears only once per article per language |
| Anchor text | Every anchor text is a real, contiguous substring of the target article's title |
| CTA density | /free-api appears at least 6 times; /register appears at least once |
| Dead links | No registration-path dead link string appears anywhere in the body |
If all nine checks pass, your article is ready for publication.
8. Summary and Next Steps
Keyword extraction and topic modeling are the highest "input-output ratio" components in a text pipeline: a few dozen lines of code can automatically surface the "highlights" from hundreds of articles. They are natural downstream applications of the Free Embedding API Complete Tutorial and the Free Semantic Textual Similarity (STS) API Complete Tutorial—vector embedding + similarity + keyword extraction form a complete "text understanding" chain; the extracted keywords can then be fed into the Free Reranker API Complete Tutorial: A Precision Filter for Your RAG Retrieval for precise recall.
Want to run this entire stack at zero cost? Start at the Free API Hub to pick a free Embedding or LLM interface as your foundation, then use the code in this article to plug in keyword extraction. If you don't have an account yet, register here for free—no credit card required, and all subsequent tutorials can be practiced hands-on immediately.
All quotas and limitations in this article were verified from Cohere / Hugging Face / KeyBERT official pages on 2026-10-02. Free quotas are subject to real-time adjustments by vendors.
7-B. How to Measure Keyword Extraction Quality
Subjective eyeballing is not enough. Use these standard metrics:
- F1 against human-labeled keywords: Have two humans independently extract top-10 keywords from 50 sample documents; treat the intersection as "gold." Compute precision, recall, and F1 of your model against that gold set. Industry baseline for KeyBERT on English news text is roughly F1 0.55-0.65.
- Downstream task lift: The most honest metric. Feed extracted keywords into your search index or tag classifier, then measure whether click-through rate, precision@k, or classification accuracy improves. If keywords do not move the downstream metric, they are decorative.
- Redundancy rate: Percentage of output keywords that are paraphrases of each other. A clean set should stay below 10 percent redundancy; enable MMR if it exceeds that.
Track these three numbers weekly during the first month of deployment. They tell you whether a model swap or a parameter change actually helped.
One final operational note: schedule a monthly review of both your keyword output and the free-tier quotas you depend on. Vendors adjust free limits frequently, and an unnoticed change from 1,000 to 100 monthly calls can silently degrade your tagging pipeline. A five-minute audit prevents days of debugging later.