How to Add a Reranker to Your RAG Pipeline (2026 Guide)
The short version
Retrieve 100, rerank, keep 10. That is the whole change. A reranker sits between your existing retrieval and your prompt assembly, needs no re-indexing, and is reversible by deleting one function call. As of October 2026 it costs roughly $0.001 per query on Voyage rerank-3 for 100 short snippets, and Voyage’s free tier gives you 200 million tokens to find out whether it helps your corpus before you spend anything. Verified October 3, 2026.
Why this works
Your embedding model encodes the query and each document separately, then compares vectors. That independence is what makes vector search fast enough to scan millions of documents — and it is also a hard ceiling on accuracy, because the model never sees the query and the document together.
A reranker is a cross-encoder: it reads the query and one candidate jointly and outputs a relevance score. That is far more accurate per document and far more expensive per document. The trick is that it only has to look at ~100 candidates instead of your whole corpus, so the cost is bounded and small.
So the two stages do different jobs: retrieval optimises recall over millions of documents; reranking optimises precision over a hundred. Running only retrieval means you are asking one model to do both, and it is bad at the second.
Step 1: Measure your baseline first
Do not add a reranker before you know whether it can help. You need two numbers from a labelled sample of 50–200 real queries:
- recall@100 — does the correct document appear anywhere in the top 100?
- precision/NDCG@10 — is it near the top?
The decision falls out immediately:
| recall@100 | NDCG@10 | Diagnosis | Action |
|---|---|---|---|
| High (>0.9) | Low | Right documents, wrong order | ✅ Reranking is the fix |
| Low (<0.8) | Low | Documents never retrieved | ❌ Fix retrieval first |
| High | High | Already working | Reranking buys little |
If recall@100 is low, a reranker cannot save you. It reorders a list; it cannot conjure documents the retriever never returned. The usual fix is hybrid search — run BM25 alongside dense retrieval and merge — because lexical and semantic search fail on different queries. Exact identifiers, error codes, product SKUs and rare proper nouns are where dense retrieval quietly loses and BM25 trivially wins.
Step 2: Widen your retrieval
Change your existing search to return 100 instead of 10. Nothing else changes.
# before
candidates = index.search(query, top_k=10)
# after
candidates = index.search(query, top_k=100)
Step 3: Add the rerank call
import voyageai
vo = voyageai.Client() # VOYAGE_API_KEY in env
def retrieve(query, k=10):
candidates = index.search(query, top_k=100)
docs = [c.text for c in candidates]
reranked = vo.rerank(
query=query,
documents=docs,
model="rerank-3-lite", # $0.02/MTok; 200M free tokens
top_k=k,
)
return [
{"text": r.document, "score": r.relevance_score}
for r in reranked.results
]
Cohere’s equivalent, if you prefer per-search billing:
import cohere
co = cohere.ClientV2()
reranked = co.rerank(
model="rerank-4-fast", # $0.002 per search
query=query,
documents=docs,
top_n=10,
)
Self-hosted, with no per-query cost:
from FlagEmbedding import FlagReranker
reranker = FlagReranker("BAAI/bge-reranker-v2-m3", use_fp16=True)
scores = reranker.compute_score([[query, d] for d in docs])
ranked = sorted(zip(docs, scores), key=lambda x: -x[1])[:10]
Step 4: Use the score, don’t just take top-10
This is the step most teams skip, and it is free accuracy. A reranker gives you a calibrated relevance score, which means you can drop weak candidates instead of always padding the prompt to exactly 10 documents.
MIN_RELEVANCE = 0.4 # tune on a labelled sample
results = [r for r in reranked.results if r.relevance_score >= MIN_RELEVANCE]
if not results:
return "I don't have information about that." # better than hallucinating
Two real benefits: the model stops being fed irrelevant context that invites hallucination, and you get an honest “I don’t know” path for out-of-scope questions. Tune MIN_RELEVANCE on labelled data — guessing it is how you silently destroy recall.
Upgrade warning: score distributions usually shift between reranker versions, which means your threshold is wrong after an upgrade and nothing errors. Voyage explicitly calibrated rerank-3’s scores to match rerank-2.5’s distribution so thresholds survive the swap; most vendors make no such promise. Re-validate your cutoff whenever you change model.
Step 5: Tune the candidate count for cost
100 candidates is a starting default, not an answer. Measure recall@k at 25, 50, 100 and 200, then pick the smallest k where recall plateaus.
recall@25 = 0.86
recall@50 = 0.94
recall@100 = 0.95 ← plateau starts at 50
recall@200 = 0.95
Here, dropping from 100 to 50 halves your reranking cost for 1% recall. Do this once and it pays forever.
The cost and latency budget
Cost, for 100 candidates averaging 200 tokens (~20K tokens per query):
| Model | Per query | Per 1M queries |
|---|---|---|
| Voyage rerank-3-lite ($0.02/MTok) | ~$0.0004 | ~$400 |
| Voyage rerank-3 ($0.05/MTok) | ~$0.001 | ~$1,000 |
| Cohere Rerank 4 Fast ($0.002/search) | $0.002 | $2,000 |
| Cohere Rerank 4 Pro ($0.0025/search) | $0.0025 | $2,500 |
| BGE-reranker-v2-m3 (self-host) | $0 + GPU | GPU only |
The billing-model trap: Cohere charges per search regardless of document length; Voyage charges per token. With 100 long documents at 8K tokens each (~800K tokens), Voyage rerank-3 costs ~$0.04 against Cohere’s flat $0.002 — a 20× inversion. Long candidates favour per-search pricing; short snippets favour per-token. Check your candidate length distribution before picking a vendor.
Latency: budget 50–300ms for 100 candidates. Self-hosted BGE-reranker-v2-m3 hits sub-40ms p50 on modest GPUs. Listwise models (Jina Reranker v3.5) score candidates concurrently and are faster on long documents. Reranking is strictly sequential after retrieval, so it adds to your critical path — if you are at a 500ms budget, measure before committing.
Common mistakes
- Reranking 10 candidates. Pointless — you keep all 10 anyway. The gain comes from having more to discard.
- Reranking before fixing recall. Precision tooling cannot fix a recall problem.
- Ignoring the score. Always returning exactly 10 documents feeds junk to the model on narrow queries.
- Reranking full documents instead of chunks when your model’s context is 32K and your documents are 100K. Chunk first, or use Jina v3.5’s 131K window.
- Not re-tuning thresholds after a model upgrade. Silent recall changes, no errors in your logs.
- Using a frontier LLM as the reranker. Orders of magnitude more expensive, seconds slower, no calibrated score.
Related: best reranker models 2026 ranked, Voyage rerank-3 vs Cohere Rerank 4 vs Jina v3.5, best embedding models 2026, best RAG frameworks 2026.
Last verified: October 3, 2026. Prices from vendor pricing pages.