← Back to articles
Tutorials

Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02)

Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02)

Semantic Textual Similarity answers a question that is deceptively simple and relentlessly frequent: given two pieces of text, how alike are they? This guide covers six routes that still work for free in 2026, a threshold calibration method you can copy directly into your own project, and seven failure modes that turn a similarity score from "plausible" into "trustworthy". Every quota below was checked against the vendor's public pricing or documentation page on 2026-10-02.

1. Why STS Is the Most Referenced, Least Explained Capability on the Site

The Free Embedding API Complete Tutorial explains how to turn text into vectors. The Free Reranker API Complete Tutorial explains how to reorder a candidate list. But in real projects the question that shows up most often is neither "how do I embed this" nor "how do I rank these". It is one of these three:

  1. Is this support ticket a duplicate of one of the eight thousand tickets we already have?
  2. Which of the three hundred canonical questions in our FAQ bank matches what the user just typed, closely enough that we can return the stored answer verbatim?
  3. Are these two crawled news articles about the same event, or two separate events that should stay separate?

All three share the same shape: you do not need a ranking, you need a single number between zero and one and a threshold to compare it against. That is exactly the domain where STS lives. Across the one hundred twenty eight articles currently published on the site, embedding tutorials, vector database roundups and retrieval augmented generation guides thoroughly cover how vectors are produced. The step that converts vectors into a decision grade score has never been covered on its own. This article closes that gap.

It also matters commercially. Duplicate detection, FAQ routing, near duplicate content clustering, deduplication before model training, and answer selection in conversational assistants are all STS problems. Teams routinely spend money on a retrieval vendor when the actual missing piece was a properly calibrated similarity threshold.

2. Three Technical Routes: Know Which Accuracy Tier You Are Paying For

Route Mechanism Typical latency Accuracy Free availability Best fit
A. Bi encoder Each text is encoded independently into a vector, similarity is the cosine between them 30 to 120 ms per text, batchable Medium, scores run 5 to 15 percent optimistic Excellent, any free embedding model works Large scale dedup, recall, clustering
B. Cross encoder Both texts are concatenated and fed through one model that emits a relevance score 80 to 300 ms per pair, cannot be precomputed High Limited, most free tiers bill per pair with small quotas Final judgement, reranking, threshold boundary review
C. LLM scoring Ask a language model to output a similarity score or an equivalence verdict 0.5 to 3 s High but high variance Good, many free chat models exist, though token cost is real Nuanced semantics, negation, cases needing a written reason

The pragmatic production combination is route A for coarse filtering, route B for final adjudication, route C for sampled audit only. Relying on route A alone with a single hard threshold is the root cause of roughly ninety percent of "the score was 0.93 but these are completely different things" incidents. For a systematic comparison of which vector models currently offer usable free quotas, see the 2026 Free Embedding Vector Model API Panorama.

One subtlety worth stating plainly: routes A and B are not "the same measurement at different quality levels". They measure different things. A bi encoder measures closeness in a shared representation space, which is dominated by topic and vocabulary overlap. A cross encoder measures interaction between the two specific texts, which is far more sensitive to relations, negation and quantity. When your evaluation set disagrees with your production behaviour, the first question should not be "is the threshold wrong" but "did I pick the wrong measurement".

3. The Six Routes That Genuinely Cost Nothing in 2026 (Measured Quotas)

Provider Free quota (verified 2026-10-02) Key models Billing unit Notes
Cloudflare Workers AI 10,000 Neurons per day, no credit card, no expiry @cf/baai/bge-large-en-v1.5, bge-m3 Neurons consumed, driven by input token count Edge network, batch requests are the most economical pattern
Jina AI Embeddings 1,000 requests per day on the free tier jina-embeddings-v3, 8192 context, 89 languages Per request Strong multilingual behaviour, friendly to long documents
NVIDIA NIM 1,000 requests per day on the free tier nvidia/nv-embed-v2, 4096 dimensions Per request Long standing MTEB strength, 32k context suits whole document vectorisation
Hugging Face Inference Providers $0.10 in credits per month for free accounts, subject to change Any embedding or text ranking model published on the Hub Compute time, extra usage requires purchased credits Since July 2025 the hf-inference provider focuses mostly on CPU class tasks such as embeddings, text ranking and text classification, which is precisely the STS workload
Google Gemini Embedding Free tier throttled by RPM, TPM and RPD; per project limits are shown live in the AI Studio console gemini-embedding-001 Per token Strong cross lingual similarity behaviour
Local sentence-transformers Unlimited, no quota at all all-MiniLM-L6-v2, bge-m3, bge-reranker-v2-m3 Your own hardware The only realistic option above roughly one million pair comparisons per day

Two clarifications that prevent expensive surprises. First, the Hugging Face free allowance is denominated in money, not requests. Ten cents stretches to several thousand short text embedding calls when routed to CPU models, but a single accidental route to a GPU backed model can consume the entire monthly allowance, so always watch the usage breakdown on the billing page rather than assuming a request count. Second, Google no longer publishes a single universal per model daily cap on its rate limit documentation page; the effective limits are displayed per project inside the AI Studio console, so any third party figure claiming "N requests per day" must be verified against your own console before you design around it.

If these free quotas are too small for your workload, the free API directory lets you compare providers side by side by category, and registering with the OneAPI unified gateway collapses many free endpoints behind a single OpenAI compatible key with unified routing and quota accounting.

4. Core Tutorial: Producing a Similarity Score You Can Actually Trust

The following Python block is the only complete code example in this article, kept deliberately because the code itself is the teaching content: batch embedding, L2 normalisation, cosine computation and threshold filtering in one pass. Any endpoint exposing an OpenAI compatible /v1/embeddings route (Cloudflare AI Gateway, Jina, NVIDIA NIM, or a local deployment) works by swapping base_url and the model name.

import numpy as np, requests

BASE = "https://api.jina.ai/v1"        # swap for any OpenAI compatible embedding endpoint
HEAD = {"Authorization": "Bearer YOUR_KEY"}

def embed(texts):
    r = requests.post(f"{BASE}/embeddings", headers=HEAD,
                      json={"model": "jina-embeddings-v3", "input": texts}, timeout=30)
    r.raise_for_status()
    return np.array([d["embedding"] for d in r.json()["data"]], dtype=np.float32)

def l2norm(x):
    return x / (np.linalg.norm(x, axis=1, keepdims=True) + 1e-9)

def similarity(pairs, threshold=0.85):
    left  = l2norm(embed([p[0] for p in pairs]))
    right = l2norm(embed([p[1] for p in pairs]))
    sims = np.sum(left * right, axis=1)      # dot product after normalisation equals cosine similarity
    return [(a, b, float(s), s >= threshold) for (a, b), s in zip(pairs, sims)]

pairs = [
    ("Cannot log in, password rejected", "Login shows the password is incorrect"),
    ("How do I change my linked phone number", "Steps to rebind a mobile number"),
    ("Can a wrong invoice be reissued", "My parcel has not moved for six days"),
]
for a, b, s, ok in similarity(pairs):
    print(f"{s:.3f}  {'similar' if ok else 'different'}  {a[:24]} <-> {b[:24]}")

Three details are routinely skipped and each one breaks the numbers. First, L2 normalisation is mandatory: without it a dot product is not cosine similarity, and scores between texts of different lengths become incomparable, which silently corrupts every threshold you set later. Second, encode the left and right sides in the same batch call: keeping both halves inside one request keeps the vector distribution consistent, whereas splitting them across separate calls with different batching boundaries can shift scores by a couple of thousandths, which is exactly the range where your threshold lives. Third, retain the full precision score in storage: rounding to two decimals at write time destroys your ability to audit boundary cases later, and you will need those three decimal places the first time you recalibrate.

5. Setting the Threshold: Never Copy Someone Else's 0.8

There is no universal best threshold. The right cut point depends on the model, the text length distribution, the domain vocabulary, and even the ratio of mixed languages in your corpus. The only defensible method is to calibrate on your own data.

Step Concrete action Output
1. Sample Pull 200 real items from production traffic, form all pairs, sort by score, take the top 40, middle 40 and bottom 40 pairs 120 candidate pairs
2. Label Humans answer one binary question only: are these about the same thing? No scoring, no gradations Positive and negative labels
3. Sweep Slide the threshold from 0.60 to 0.98 in steps of 0.01, computing precision and recall at each point A precision recall curve
4. Choose Deduplication favours precision at or above 0.95; FAQ routing favours recall at or above 0.90; pick according to which error is more expensive A single number
5. Regress Add twenty fresh samples each week and rerun; if the optimal cut moves by more than 0.03, recalibrate and version the threshold A versioned threshold

Reference ranges, offered as a sanity check rather than a preset: bge-large-en-v1.5 class models on short customer support text typically reach 0.95 precision somewhere in the 0.88 to 0.93 band. For cross lingual comparison, the same precision target usually requires lowering the threshold by 0.04 to 0.07. Changing the model always requires recalibration, which is the single most commonly skipped step and the most common source of silent production regressions.

A practical shortcut when you have no labelled data yet: build a small anchor set of ten pairs you know are unrelated and ten you know are identical, run it once a day, and track the score distributions. Anchor drift is an early warning that the provider changed something underneath you.

6. Seven Failure Modes That Silently Distort Your Scores

Each of these has a characteristic production incident signature. Recognising the signature is faster than debugging from raw numbers.

6.1 Length asymmetry. Comparing an eight word question against a three hundred word document systematically depresses cosine similarity, because the long document's vector averages over many subtopics while the short question concentrates on one. The fix is not to change the model: chunk the long text into segments of roughly the same length as the short side, score the question against every chunk, and take the maximum. This converts an apples to oranges comparison into apples to apples, and it typically lifts true positives by 0.05 to 0.12 while barely touching true negatives.

6.2 The negation trap. "Refunds are supported" and "Refunds are not supported" routinely score above 0.95 with a bi encoder, because negation changes the truth value without changing the topic, and bi encoders are overwhelmingly topic sensitive. Any workflow where a wrong match causes an incorrect action, such as auto answering a support question, must send borderline pairs through a cross encoder or an LLM check. The Free Reranker API Complete Tutorial covers the cross encoder side of that decision in depth.

6.3 Numbers and identifiers. "Order 88213" and "Order 88214" look nearly identical to an embedding model and are semantically unrelated. Extract numbers, IDs, product codes and currency amounts with regular expressions before similarity scoring, compare them exactly, and reject the pair outright on mismatch. Vectors are for recall, exact matching is for correctness; conflating the two is how a dedup pipeline silently merges different orders.

6.4 Score baseline inflation. Different models have different baseline behaviours, and even a single model shifts its baseline between versions. Maintain an anchor set of ten known dissimilar pairs, run it with every batch, and watch the mean anchor score. If your anchors used to average 0.31 and now average 0.44, your hard coded threshold no longer means what it meant last quarter.

6.5 Silent batch truncation. When you submit two hundred texts in one request, some free endpoints quietly cap the batch and return fewer vectors than you sent, without an error code. Assert that the returned vector count equals the input count on every call, and split and retry when it does not. This bug is invisible until someone notices that a whole shard of documents was never indexed.

6.6 Cache misses from whitespace noise. The same sentence with a trailing space, a full width character, or a different newline convention is billed as a new text and embedded as a new vector. Normalise before hashing: collapse whitespace, fold full width punctuation to half width, unify unicode forms, and optionally strip trailing punctuation. Normalisation alone commonly lifts cache hit rate by fifteen to twenty five percentage points, which converts directly into quota saved.

6.7 Mixed language tokenizer bias. In code switching corpora, English tokens frequently dominate the vector even when the informational content is in another language, so mixed sentences cluster with the wrong group. Route any text above roughly thirty percent foreign language content through its own threshold, or through a model that was explicitly trained for multilingual alignment.

7. Production Engineering: Batching, Caching, Backoff and Fallback

Once the algorithm is settled, the real cost of running STS at scale is quota management, not mathematics. Four practices carry almost all of the savings.

  1. Local vector cache. Key the cache by the hash of the normalised text and store the vector. Hit rates of sixty to eighty five percent are typical for support and FAQ traffic, which multiplies your effective free quota by the same factor.
  2. Batch consolidation. Merge fifty to two hundred texts into a single embedding call. For request counted tiers such as the 1,000 per day caps, this is the difference between unusable and comfortably sufficient, because request count drops by the batch size rather than by one.
  3. Exponential backoff with jitter on 429. On a rate limit response, wait one second, then two, then four, adding random jitter so that a fleet of workers does not retry in lockstep and trigger a second wave of throttling. Without jitter, retries cluster and the tail latency looks like an outage when it is only a stampede.
  4. Fallback chain across providers. When the primary quota is exhausted, fail over to a secondary provider, for example Jina to Cloudflare to a local bge-m3, and record the model name in the result metadata so you can detect threshold drift caused by the switch. A failover that changes the score distribution without changing your threshold is worse than an outage, because it produces confidently wrong answers.

The complete backoff, budget and fallback implementation is documented in Free API Cost and Quota Control in Practice; the pattern above only lists the parts that specifically matter for embedding workloads. If your similarity scores feed a retrieval pipeline rather than a dedup job, the storage and top k selection layer is covered by the Free Vector Database API Power Rankings, and the end to end wiring from documents through embeddings to answers is walked through in Building a RAG Knowledge Base with Free Embedding APIs.

8. How to Measure Whether Your STS Pipeline Is Working

A similarity pipeline without metrics is a guess with an API key attached. Four measurements are worth automating in the first week.

Metric Definition Healthy target What it tells you
Precision at threshold Of pairs flagged similar, the fraction actually similar 0.95 or higher for dedup Cost of false positives: wrongly merged records
Recall at threshold Of truly similar pairs, the fraction flagged 0.90 or higher for FAQ routing Cost of false negatives: missed duplicates, unanswered questions
Boundary volume Fraction of pairs landing within plus or minus 0.03 of the threshold Under 10 percent If high, your threshold sits in the middle of the distribution and needs a better model or a two stage design
Anchor drift Change in mean anchor pair score versus last week Under 0.02 Provider or model changed underneath you without warning

Interpretation guidance: when boundary volume is high, no amount of threshold tuning will help, because the score is not separating your classes. That is the signal to add a cross encoder second stage or to switch model families. When precision is fine but recall is poor, your threshold is simply too strict. When both are poor, the model is wrong for the domain and calibration is wasted effort, which is why step one of the calibration table is labelling your own data before touching any number.

9. Frequently Asked Questions

Can I use one of the free chat models instead of a dedicated similarity endpoint? Yes, but treat it as route C: sampled audit, not bulk scoring. Chat models are slow, expensive in tokens, and their numeric scores are not calibrated, meaning 0.9 from one prompt today and 0.9 from a rephrased prompt tomorrow are not comparable quantities. Use them where you need a written reason, not where you need a sortable number.

How many pairs does 10,000 Cloudflare Neurons actually buy? Neuron consumption scales with input tokens, so the honest answer depends on your text length; measure it on your own corpus with a few hundred calls and read the usage dashboard, then extrapolate. Anyone quoting a fixed pair count is quoting a different workload than yours.

Is a threshold of 0.85 safe as a starting point? Only as a placeholder until you calibrate. It is fine for a demo and negligent for production, because the correct number for your data might be 0.72 or 0.94 depending on the model and the domain.

Should I dedup with cosine similarity or with exact matching? Exact matching for identifiers, amounts and code; similarity for natural language paraphrase. Real systems need both, applied in the order that is cheapest first: normalise, exact match, then similarity.

Does this work for Chinese and English at the same time? It does with multilingual models such as bge-m3, but expect the cross lingual threshold to be a few points lower than the monolingual one, and always calibrate the mixed case separately.

10. Your Action Checklist

  1. Pick the route: under roughly ten thousand pair comparisons a day, start with Cloudflare or Jina free tiers; above that, deploy bge-m3 locally.
  2. Build the evaluation set: 120 pairs, binary human labels. This step is non negotiable.
  3. Calibrate: run the sweep in section 5 and produce your own threshold, never a copied one.
  4. Add cache and batching first: reduce request count before you argue that quotas are too small.
  5. Instrument: log anchor scores daily, boundary volume per model, and per provider consumption.
  6. Define the review policy: pairs within 0.03 of the threshold go to a cross encoder or LLM second opinion.

Follow those six steps and you end up with a zero cost, explainable, regressable text similarity pipeline. Live availability and quota changes for all six routes are tracked in the free API directory; to consolidate these endpoints behind one key, one quota and one bill, register with the OneAPI gateway and switch upstreams using a single OpenAI compatible protocol.

11. Throughput and Latency: What to Measure Before You Commit

Free tier selection is not only about quota size; two providers with identical request caps can differ by an order of magnitude in usable throughput. Four numbers decide whether a route survives contact with your workload.

Measurement How to collect it Decision it drives
Cold start latency First call after ten minutes idle, repeated five times, report the median and the worst Whether a user facing feature can tolerate the provider at all
Warm per text latency Median of a five hundred text batch after two warmup calls Whether you can score a queue in real time or must batch overnight
Batch scaling curve Latency at batch sizes 1, 8, 32, 128 on identical text length The exact batch size where further gains flatten, which is your optimal request shape
Sustained throttle point Ramp concurrency until the first 429, record requests per minute at that moment Your real ceiling, which is almost always lower than the advertised cap

The batch scaling curve deserves special attention because it is where free tiers differ most. Some endpoints pay a fixed per request overhead, meaning batch size thirty two costs barely more than batch size one and you should batch aggressively. Others bill linearly per item, where batching only saves network round trips. Cloudflare's neuron based accounting falls in the second family, so the lever there is shortening input text through normalisation and chunk selection rather than growing batches; the full per provider breakdown is maintained in the free API directory.

Cross regional latency is the second silent differentiator. An endpoint that measures 40 ms from a test machine in one region can measure 400 ms in another, and similarity scoring is latency sensitive in any interactive path because you typically score dozens of candidates per user action. Measure from the region your users are actually in, not from a benchmark table, and treat any published figure as a hypothesis until your own probe confirms it. Category level availability notes for embedding and similarity services, including which ones currently respond from which regions, are tracked in the free API directory.

Finally, record the score distribution, not just the average. A model that produces a tight bimodal distribution on your data, with clear separation between related and unrelated pairs, is worth far more than a model with a marginally better public benchmark and a smeared distribution, because the first gives you a threshold and the second gives you a coin toss with extra steps. Plotting that histogram takes one line of code and is the highest value diagnostic available before you commit to a provider; the free API directory records which of the six routes above currently return usable vectors for a given text length, which is the fastest way to shortlist candidates before you spend an afternoon benchmarking them yourself, and it is the same diagnostic you rerun whenever you change models or the anchor drift alarm fires.

12. Where to Look Next

This article covers the scoring step. The surrounding pipeline is documented in depth across the site, and the current live status of every provider mentioned here, including quota changes and endpoint availability, is maintained centrally in the free API directory. Teams that need one credential, one rate limit view and one bill across all of these providers can set that up by registering with the OneAPI gateway and pointing a single OpenAI compatible client at it.

More in this category

Free Text Summarization API Complete Tutorial: Let LLMs Compress 1M-Word Documents into 100 WordsFree Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)Free Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)Free Time Series Forecasting API Complete Tutorial: Zero-Cost “Crystal Ball” for Sales/Inventory/Energy Prices (Verified 2026-10-03)Free Voice Cloning API Complete Tutorial: Clone Your Signature Voice from a Reference Clip

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.