AI agents · OpenClaw · self-hosting · automation

Quick Answer

Best Reranker Models 2026: Ranked by Quality and Price

Published:

The ranking (October 2026)

This is a recommendation order, not a single score ordering — the right reranker depends mostly on billing model, deployment constraint and candidate length. Prices verified October 3, 2026.

#ModelPriceContextDeployPick it when
1Voyage rerank-3-lite$0.02 / MTok32KAPIDefault. Best value, 200M free tokens
2Voyage rerank-3$0.05 / MTok32KAPIMax quality, long documents, code
3Cohere Rerank 4 Fast$0.002 / search32KAPIFlat cost, long candidates, 100+ languages
4BGE-reranker-v2-m3Free (Apache 2.0)~8KSelf-hostZero per-query cost, data residency
5Jina Reranker v3.5Token packages131KAPI + self-hostHuge candidates, air-gapped, low latency
6Cohere Rerank 4 Pro$0.0025 / search32KAPIEnterprise support, max Cohere accuracy
7Qwen3-Reranker 8BFree (open weights)32KSelf-hostOpen weights with more headroom, multilingual
8mxbai-rerank-large-v2Free (Apache 2.0)~8KSelf-hostLatency-sensitive self-hosted reranking

1. Voyage rerank-3-lite — the default

$0.02 per million tokens, 32K context, 200 million free tokens per account.

Voyage states that rerank-3-lite matches the retrieval quality of the previous-generation rerank-2.5 at 40% of the price, and beats Cohere Rerank v4.0 Pro by 2.23% NDCG@10 across Voyage’s 95-dataset suite. The 200-million-token free grant means you can validate reranking on your actual corpus for literally nothing before committing budget. Start here.

2. Voyage rerank-3 — the quality pick

$0.05 per million tokens, 32K context, instruction-following.

Released September 30, 2026. On Voyage’s suite it is the top performer on every first-stage retriever tested (BM25, OpenAI text-embedding-3-large, voyage-3-large, voyage-4-large), beating Cohere Rerank v4.0 Pro by 2.72% overall and by 13.84% on long documents (LongEmbed: NarrativeQA, SummScreenFD, QMSum). Code retrieval improved 2.17% over rerank-2.5 — Voyage explicitly retrained for coding-agent queries, which now make up a large share of its traffic.

Two practical virtues beyond score: it is a drop-in replacement for rerank-2.5 (same API, same price, same context), and relevance scores are calibrated to match rerank-2.5’s distribution, so existing score thresholds keep working. That is the difference between a one-line swap and a week of re-tuning.

Caveat: these are vendor-published numbers, with Voyage’s own embedder as the headline first stage and Qwen3-Reranker-8B’s cost assumed rather than measured. A 2.72% margin is real but modest. Test the 13.84% long-document claim yourself if long candidates are your workload — that is the one that would change an architecture.

3. Cohere Rerank 4 Fast — the flat-rate choice

$0.002 per search, 32K context, 100+ languages.

Cohere bills per search — one query plus up to 100 documents, regardless of document length. This inverts the cost maths against token-billed rivals. For 100 long documents at 8K tokens each (~800K tokens), Voyage rerank-3 costs about $0.04 while Cohere charges a flat $0.002, roughly 20× cheaper. For 100 short 200-token snippets, Voyage wins. Know your candidate length distribution before choosing a billing model — this is the single biggest cost lever in reranking and almost nobody checks it.

Rerank 4’s 32K context (quadrupled from v3.5’s 4,096 tokens) is what makes long-candidate reranking viable at all on Cohere. Rerank 4 Pro at $0.0025 buys maximum accuracy and deeper reasoning; Rerank v3.5 remains at $0.001 if 4K context is enough.

4. BGE-reranker-v2-m3 — the free default

Apache 2.0, ~568M parameters, multilingual, single consumer GPU.

Still the most widely deployed open-weight reranker in production RAG, and the correct answer when per-query fees or data residency are hard constraints. Sub-40ms p50 latency is achievable on modest hardware. The trade-off is out-of-domain generalisation: expect to fine-tune on your domain to match proprietary quality, and budget engineering time accordingly. Free per query is not free in total.

5. Jina Reranker v3.5 — the long-context and self-host option

0.6B parameters, 131K context, listwise, released July 27, 2026.

The standout spec is the 131K context window — 4× Voyage and Cohere — which matters when you rerank genuinely large documents rather than chunks. Its listwise architecture scores multiple candidates concurrently instead of pair-by-pair, which is where its latency advantage comes from (Jina reports up to 56% faster than v3 on long documents). Reported strength in legal and structured-data retrieval. Available via the Jina API, Hugging Face weights, Elastic’s inference integration, and air-gapped deployment. Check the commercial licence for production use of the weights.

6–8. The rest, briefly

Cohere Rerank 4 Pro ($0.0025/search) — the accuracy-first Cohere tier, with the strongest enterprise support story in the category. Qwen3-Reranker (0.6B / 4B / 8B, open weights) — strong multilingual accuracy across 100+ languages, 32K context on the 8B, a credible open alternative when BGE is not enough. mxbai-rerank-large-v2 (Apache 2.0) — RL-tuned, permissively licensed, a good latency-sensitive self-hosted pick.

The decision in four questions

  1. Can you call an external API? No → BGE-reranker-v2-m3, Jina v3.5 self-hosted, or Qwen3-Reranker.
  2. Are your candidates long (>2K tokens each)? Yes → Cohere’s per-search billing, or Jina v3.5 for very long candidates. No → Voyage’s per-token billing is cheaper.
  3. Do you need 100+ languages with vendor support? Yes → Cohere Rerank 4.
  4. Otherwise → Voyage rerank-3-lite, upgrade to rerank-3 if evaluation shows the quality is worth 2.5×.

What not to do

Do not use a frontier LLM as your reranker. Asking a $2/$10-per-MTok model to rank 100 candidates costs orders of magnitude more than a $0.05/MTok dedicated reranker, adds seconds of latency, and returns no calibrated score you can threshold on. Dedicated rerankers exist because this task is narrow enough to solve at 0.6B parameters. Reserve LLM ranking for when you need the ranking explained in prose.

Do not skip the baseline measurement. Reranking helps most when your first-stage retrieval has decent recall but poor precision at top-10. If recall@100 is already bad, a reranker cannot invent documents the retriever never returned — fix retrieval first.

Related: Voyage rerank-3 vs Cohere Rerank 4 vs Jina v3.5, how to add a reranker to your RAG pipeline, best embedding models 2026, best RAG frameworks 2026.

Last verified: October 3, 2026. Prices from vendor pricing pages; comparative scores are Voyage AI’s published benchmark results.

Sources