โ† Back to articles
Tutorials

Free Reranker API Complete Tutorial: A Precision Filter for Your RAG Retrieval

Free Reranker API Complete Tutorial: A Precision Filter for Your RAG Retrieval

RAG (Retrieval-Augmented Generation) has become the most popular assembly pattern in the free API ecosystem. You start by turning text into vectors with a Free Embedding API, then hand the vectors to a Free Vector Database for semantic recall. But many builders hit the same experience: in the Top-10 from vector search, the first two results are always right, while the last eight are full of noise that looks relevant but is actually irrelevant. That is exactly when you need a reranking layer.

This article covers three things systematically: what shortcoming of vector search reranking actually fixes, which Reranker APIs are genuinely free to call in 2026, and how to raise retrieval precision without spending a fortune. Every model name and free-tier quota below is confirmed against the vendor's public page โ€” nothing is fabricated.

1. Why embedding retrieval is "fast but not precise enough"

Vector retrieval is built on a bi-encoder: the query and each document are encoded independently into vectors, then compared by cosine similarity. This architecture wins on speed โ€” it can recall a Top-50 or even Top-100 across millions of items in milliseconds.

But the bi-encoder pays a structural price: the query and the document never meet each other during encoding. The model "reads" the query and the document separately, but never compares them word by word. As a result, "AAPL stock price" and "Apple iPhone price" can end up close in vector space โ€” they share too many literal terms and directional signals, yet the user's true intent (finance vs. consumer electronics) is completely different.

This is the root cause of vector search's "high recall, low precision". It excels at the floor (fetching everything relevant), but not at precise ranking (putting the most relevant items first).

2. What reranking fixes: bringing in a model that has "met" both sides

A reranker takes a different architectural path: the cross-encoder. It feeds the (query, document) pair into the model together as one input, lets the model understand the interaction between them token by token, then outputs a single relevance score.

Because the query and document truly face each other, the cross-encoder captures details the bi-encoder cannot โ€” pronoun references, negations, causal relationships, domain-specific synonyms. The cost is speed: it cannot exhaustively compare against the whole corpus, so it can only re-rank the top 50 or 100 already recalled by vector search.

A standard RAG pipeline is therefore two-stage:

  1. Broad recall: use the bi-encoder embedding to recall Top-50~100 from the whole corpus (fast, floor);
  2. Fine reranking: use the reranker to score those dozens again and output only the Top-5~10 to the LLM for generation (precise, refined).

Each stage does its own job, which is how you get both "fast" and "precise" at once.

3. The full landscape of free Reranker APIs in 2026

Below I organize the landscape along two lines: managed APIs (ready out of the box) and open weights (free to self-host). Model names and quotas are verified against vendor public pages.

3.1 Managed APIs (register and go)

Vendor Model Free tier Highlights License
Jina AI jina-reranker-v3.5 100 RPM / 100K TPM 0.6B listwise, 131K context, 100+ languages CC-BY-NC-4.0 (weights); commercial via API
Cohere Rerank 4 Fast / Pro Limited trial key Multilingual managed, up to 1000 docs per search Closed
Voyage AI rerank-2.5 / lite Limited trial Enterprise search, bundled with embeddings Closed (acquired by MongoDB)
mixedbread mxbai-rerank-large-v2 API metered per token Apache 2.0, open weights + managed dual track Apache 2.0

3.2 Open weights (free to self-host, GPU cost on you)

Model Params Context License Notes
BGE Reranker v2-m3 (BAAI) 568M 8192 Apache 2.0 Multilingual cross-encoder; de facto standard for Chinese
Qwen3 Reranker 0.6B/4B/8B 32K Apache 2.0 Alibaba family, strong Chinese-English benchmarks
mxbai-rerank-large-v2 ~2B โ€” Apache 2.0 Second generation, multilingual
ColBERTv2 โ€” โ€” MIT (Stanford) Late-interaction route

The key distinction: BGE v2-m3, Qwen3, mxbai, and ColBERT are all commercially usable open source; Jina's weights are CC-BY-NC-4.0 (non-commercial), so commercial use requires its hosted API or a commercial license from Elastic โ€” this is the license trap beginners stumble into most often.

There is enough free headroom to get started on the free API channel โ€” both tracks have accessible tutorials and tested numbers.

4. Why the free tier is enough to get you started

Many people misunderstand rerankers, assuming they eat as much compute as LLMs and therefore must cost money. In reality, cross-encoders are generally small โ€” BGE v2-m3 has only 568M parameters and Jina v3.5 only 0.6B, far smaller than the 70B-class chat models. That means two things.

First, free-tier quota is enough for trial runs. Take Jina's Reranker API: the free key gives 100 RPM / 100K TPM. One rerank request carries dozens of documents at a few hundred tokens each, so even a developer repeatedly debugging locally and running a full small RAG eval set will struggle to hit the wall in a single day.

Second, the self-hosting hardware bar is lower than you think. If you go the BGE v2-m3 route, a 568M-parameter model, once quantized, can even run on a consumer GPU or some CPUs, making it a zero-API-bill option for personal projects and small teams.

So the reranking layer has almost no "you must pay before you can even start" barrier โ€” it is actually the cheapest, fastest-returning link in the whole RAG stack.

5. Listwise vs. pointwise: a small but critical classification

Rerankers split internally into two camps:

  • Pointwise: each document is scored against the query independently, without affecting each other. Representatives are BGE v2-m3 and early Jina v2.
  • Listwise: the query and all candidate documents share one context window, and a single forward pass compares all documents against each other before scoring. Representatives are Jina v3 / v3.5.

The intuitive advantage of listwise is that "relevance is relative" โ€” how high a document should rank often depends on what else is in the candidate set. This gives larger gains on contexts requiring cross-document comparison such as structured and legal text, but it also requires a longer context window (Jina v3.5 supports up to 131K tokens).

6. A free, hands-on example: wiring reranking into your pipeline

Below is a minimal example demonstrating the two-stage "vector recall + rerank". This is the only core teaching content kept as code; the rest is explained in prose.

# Stage 1: broad recall via vector search
# Assume docs are already indexed; use embeddings to fetch Top-50
recalled = vector_search(query="AAPL stock price today", top_k=50)

# Stage 2: reranker re-scores the survivors
import requests

resp = requests.post(
    "https://api.jina.ai/v1/rerank",
    headers={"Authorization": "Bearer YOUR_KEY"},
    json={
        "model": "jina-reranker-v3.5",
        "query": "AAPL stock price today",
        "top_n": 5,
        "documents": [d["text"] for d in recalled],
    },
)
top5 = resp.json()["results"]  # ordered by relevance, high to low

This snippet reveals the core contract of the two-stage design: the reranker does not search the whole corpus itself, it only finely re-orders the candidate set fed from upstream. So the quality of upstream recall (which Embedding you use, how you chunk) directly determines the ceiling of reranking. To understand recall-layer selection systematically, revisit the Embedding and vector-database topics on the free API channel.

7. A selection decision framework

Your constraint First pick Why
Chinese RAG, zero API bill, can self-host BGE Reranker v2-m3 Apache 2.0, de facto Chinese standard
Want listwise, 131K context, multilingual Jina v3.5 managed API Free tier starts at 100 RPM
Already on Voyage embeddings Voyage rerank-2.5 Single vendor, unified billing
Want open + managed dual track, clean license mxbai-rerank-large-v2 Apache 2.0
Long documents, need token-level relevance ColBERTv2 Late-interaction route

One-line principle: for ordinary Chinese RAG, self-hosting BGE v2-m3 costs the least; if you want zero-ops and fast time to production, Jina's free tier is currently the lowest-friction managed entry point.

8. Four frequent pitfalls

  1. License: Jina's weights are CC-BY-NC-4.0 (non-commercial); for commercial use you must go through the API or a commercial license โ€” don't ship the weights directly into a paid product.
  2. Using public benchmarks as the only basis: BEIR/MTEB/MIRACL are a starting point; real corpora have different term distributions, so always re-test on your own 100โ€“500 query golden set.
  3. Ignoring the latency budget: rerank latency stacks on top of recall, and managed APIs add network round-trip; measure p50/p95 before production.
  4. Running an English-only reranker on Chinese corpora: ColBERTv2 base and most ms-marco family are English-only; running them on Chinese loses precision โ€” check the language coverage on the model card first.

8.5 How to evaluate a reranker properly

Choosing a reranker on marketing claims alone is a recipe for disappointment. The only defensible way is to measure it on your own workload. Here is the full evaluation recipe.

Step 1 โ€” Build a golden set. Pull 100 to 500 real queries from your production logs. For each query, label the top-10 candidate documents by relevance. Commit the labels to a versioned dataset. Without a golden set, every downstream number is guesswork.

Step 2 โ€” Run two passes. Retrieve Top-50 with your bi-encoder once, then re-rank that same Top-50 with each candidate reranker. Compute nDCG@10, MRR@10, and Recall@10 per model. Do not re-run retrieval separately per model โ€” you want to isolate the reranker's contribution, not the embedding's.

Step 3 โ€” Cost-adjust. Record end-to-end latency, GPU memory (for self-host), and cost per query for each candidate. A model that wins on accuracy but blows the latency budget is not a win in production. The order of priorities should be: accuracy first, then latency, then cost โ€” and only compromise in that order.

Step 4 โ€” Gate the change. Add the reranker eval to CI. Any future swap of the reranker, the embedding model, or the chunking strategy should re-run the eval and surface a regression before it merges. This one habit prevents the classic "we upgraded the model and the answers got worse" regression.

The four metrics matter in different ways: nDCG@10 rewards ranking the most relevant document at the top; MRR@10 rewards where the first relevant document lands; Recall@10 checks whether the relevant documents survived into the final ten at all. Track all three โ€” they measure different failure modes.

8.6 Latency reality check: what "fast" actually means

Reranker latency is not a single number, and vendor marketing rarely tells you the full story. It scales with three variables: document count, document length, and batch size.

As a concrete reference, Jina publishes a latency table for reranking one query against 100 documents. With 64-token documents and a 64-token query, the round trip is roughly 150 milliseconds. Push each document to 4096 tokens and the same operation takes about 3.5 seconds. Double the query to 512 tokens and it stretches toward 7 seconds. The pattern is clear: latency is dominated by the total token volume โ€” the sum of query plus all documents โ€” not by the raw number of documents.

Three implications follow. First, chunking discipline matters for latency, not just accuracy โ€” oversized chunks silently bloat every rerank call. Second, managed APIs add network round-trip on top of model inference, so measure from your own region, not from the vendor's published latency. Third, if your end-to-end RAG budget is under 200 milliseconds total, a cross-encoder reranker may not fit at all โ€” in that case the bi-encoder must already hit your recall target on its own.

The practical takeaway: measure p50 and p95 on your real document-length distribution and your real hardware. A quoted "150 ms" means nothing until you reproduce it against your own chunks.

8.7 A compact runbook

Before you ship a reranker into production, walk this checklist top to bottom.

  1. Confirm the license fits your use case โ€” commercial vs. research, hosted vs. self-host.
  2. Confirm the language coverage matches your corpus โ€” don't run an English-only model on Chinese text.
  3. Build a 100โ€“500 query golden set, labeled by hand.
  4. Measure nDCG@10, MRR@10, and Recall@10 for at least two candidates.
  5. Record latency (p50/p95) and cost per query on your real chunks.
  6. Add the eval to CI so future swaps surface regressions automatically.
  7. If on a managed API, verify the free-tier RPM/TPM covers your trial volume before scaling.

8.8 When you should skip the reranker entirely

Reranking is powerful but not universal. There are four situations where you should skip it and keep the pipeline lean.

First, when the end-to-end latency budget is below roughly 200 milliseconds and the bi-encoder already hits the recall target โ€” the cross-encoder simply will not fit. Second, when the corpus is small enough that Top-3 retrieval is already precise; a reranker adds a step and a dependency for negligible gain. Third, when the downstream LLM genuinely tolerates noisy context โ€” rare in practice, but some summarization or broad-topic tasks are robust to a few off-topic chunks. Fourth, when your own eval shows no nDCG lift on the real workload; trust the measurement, not the hype.

The skip case is the exception, not the rule. Most production RAG pipelines benefit from a reranker โ€” but every decision should be driven by a number you measured yourself, not by a benchmark someone else ran on a corpus you have never seen.

8.9 Architectures compared: cross-encoder, late interaction, listwise

The three reranking architectures trade accuracy against cost in different ways, and understanding the trade-off tells you which one fits your constraint.

Cross-encoder is the classic. Query and document are concatenated into one input and pass through the model jointly, producing a single relevance score. It is the most accurate per pair because every token of the query can attend to every token of the document. The cost is that scoring N candidates requires N forward passes โ€” you cannot precompute anything. This is why cross-encoders live after recall, never instead of it. BGE v2-m3 and Jina v2 belong here.

Late interaction (ColBERT family) is a hybrid. Documents are pre-encoded into per-token embeddings and stored; at query time, the model computes a similarity matrix between query tokens and document tokens and takes a max-sim per query token. Because document vectors are precomputed, late interaction is far cheaper at query time than a full cross-encoder pass, yet it retains token-level granularity that a bi-encoder loses. The catch is storage: per-token embeddings inflate the index roughly 30 to 100 times over a sentence-vector index, though PLAID-style compression brings it back to a working size. Choose late interaction when your corpus is long, your queries are latency-sensitive, and you can afford the index.

Listwise is the newest of the three. The query and all candidates share one context window and are scored in a single forward pass, so the model compares documents against each other before ranking. Relevance is treated as relative rather than absolute, which matches how a human actually ranks a result page. This is the architecture behind Jina v3 / v3.5, and it explains why those models want enormous context windows (131K tokens) โ€” they must hold the query plus dozens of documents simultaneously.

In practice: cross-encoder is the default for most teams; late interaction earns its complexity on long-document corpora with tight query latency; listwise is worth trying when candidate-set-relative ranking matters, such as legal or structured-data retrieval.

8.10 Wiring a reranker into LangChain and LlamaIndex

If you already run a RAG framework, adding a reranker is usually a ten-line change rather than a rewrite.

In LangChain, the pattern is a contextual compressor: the retriever returns Top-50, the compressor reranks and squeezes it to Top-5, and the compressed documents go to the prompt. Cohere ships an official CohereRerank compressor; Jina's API is simple enough to wrap in a custom compressor with an HTTP call. The one contract to respect: the compressor must return documents in the new relevance order, not just new scores, or the downstream prompt will still see the old ranking.

In LlamaIndex, the reranker sits as a node_postprocessor between retrieval and response synthesis. Both CohereRerank and SentenceTransformerRerank (for self-hosted BGE) exist as first-class postprocessors. The configuration knob that matters most is top_n โ€” set it to the number of chunks your prompt budget can actually hold, typically 3 to 8 for a 4K to 8K context window.

For either framework, the same two-stage contract applies: the reranker only reorders what the retriever hands it. If retrieval misses the right document entirely, no reranker can recover it. Fix recall problems upstream (chunking, embedding choice) and precision problems downstream (reranking, prompt assembly). The free API channel has companion tutorials for both halves of that pipeline.

8.11 Free-tier economics and the upgrade path

The free tiers described above are genuinely usable, not demo-only, but they differ in what they throttle.

Jina throttles by rate: 100 requests per minute and 100,000 tokens per minute on the free key. For a personal knowledge base or an evaluation harness, that ceiling is effectively invisible โ€” you would need to fire dozens of rerank calls per second to notice it. The throttle that bites first is TPM, because a single request with 50 long documents can burn tens of thousands of tokens by itself.

Cohere and Voyage throttle by trial credits: a limited count of requests tied to the trial key, after which you move to a paid plan. That is fine for evaluation and small pilots, but you should budget before production traffic.

Self-hosted open weights have no request ceiling at all โ€” the only cost is the GPU you already own. A 568M-parameter BGE v2-m3 on a mid-range consumer card can serve a small team comfortably, and quantization buys further headroom.

A sensible upgrade path: prototype on Jina's free tier (zero setup), benchmark two or three candidates on your golden set, and if the winner is an Apache 2.0 model and traffic grows, migrate to self-hosting for zero marginal cost. The full catalog of free endpoints, including every reranker listed here, is indexed on the free API channel. Because the interface is identical (query in, ranked documents out), the migration is a config change, not a code rewrite.

8.12 Error handling you will actually need

Three failure modes show up constantly in production reranking.

429 rate-limit responses. Managed free tiers will throttle you during bursts โ€” evaluation runs, batch indexing jobs, traffic spikes. The fix is the same pattern used for every free API: exponential backoff with jitter, a hard retry cap, and a fallback that skips reranking entirely and serves the raw bi-encoder order. A degraded but responsive answer beats a correct answer that arrives three retries late. The full playbook for 429 handling across multi-channel setups is covered in our companion guide on Cost and Quota Control.

Oversized payloads. A rerank request carrying 50 documents of 4,000 tokens each can exceed context or TPM limits depending on the model. Defend by capping per-document length before submission (truncate to the first N tokens or use your chunk boundaries), and by trimming candidate count โ€” reranking Top-25 usually captures most of the lift of reranking Top-100 at a quarter of the cost.

Silent ordering bugs. The most dangerous failure is not an error but a wrong pipeline: your code calls the reranker, receives the results, and then โ€” through a mapping bug โ€” passes the original document order to the prompt anyway. Unit-test for this explicitly: craft a query where the reranker must swap two documents, and assert the swapped order reaches the generator.

8.13 The 2026 timeline in brief

For context on how fast this space moved: BGE v2-m3 established the Apache 2.0 multilingual baseline in 2024; Cohere shipped Rerank 3.5 in December 2024 and its Rerank 4 generation (Fast and Pro) across 2025 to 2026; Jina moved from v2 (pointwise, 2024) to v3 and then v3.5 (listwise, 131K context) in 2025 to 2026; Voyage launched rerank-2.5 as its flagship while being absorbed into MongoDB; Qwen3 Reranker brought a 0.6B-to-8B Apache 2.0 family mid-2025.

The direction is unmistakable: open-license multilingual rerankers now match or beat closed managed models on standard benchmarks, while managed APIs compete on context length and latency. For a builder assembling a zero-cost stack, that means the free options are no longer compromises โ€” they are first-class choices.

9. Conclusion

Reranking is the highest-ROI layer in the RAG pipeline โ€” it does not ask you to replace your existing Embeddings and vector database; it simply inserts a cross-encoder between recall and generation to lift Top-10 precision by a notch. The 2026 free ecosystem is mature enough: for Chinese self-hosting there is Apache 2.0 BGE v2-m3, and for zero-ops managed there is Jina v3.5 starting at a 100 RPM free tier.

Once you bolt on this layer, your RAG Knowledge Base truly evolves from "can search" to "searches precisely". The broader lesson generalizes beyond reranking: in a free API stack, each layer has a specialized job, and the wins come from composing them correctly rather than from any single component. The embedding layer owns breadth, the vector database owns storage and speed, the reranker owns precision, and the LLM owns synthesis. When retrieval quality disappoints, the first diagnostic question should always be "which layer is failing" โ€” recall failure is an embedding or chunking problem, precision failure is a reranking problem, and answer failure is a prompt or model problem. Treating them as one undifferentiated "search quality" issue is the single most common mistake in DIY RAG. For more free API integrations and selection, keep browsing the latest articles on the free API channel, covering vector retrieval, speech recognition, multimodal understanding and the whole chain. All these capabilities are reachable through a single OneAPI Unified Gateway with one registration and shared models across vendors โ€” sign up here: register for free.

Further reading

More in this category

Free Text Summarization API Complete Tutorial: Let LLMs Compress 1M-Word Documents into 100 WordsFree Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)Free Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)Free Time Series Forecasting API Complete Tutorial: Zero-Cost โ€œCrystal Ballโ€ for Sales/Inventory/Energy Prices (Verified 2026-10-03)Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide โ€” sign up and get bonus credits.