Voyage Rerank 3 vs Cohere Rerank 4 vs Jina v3.5 (2026)
The short answer
Voyage rerank-3 (September 30, 2026) is the quality leader on its own benchmark; Cohere Rerank 4 is the safest managed API; Jina Reranker v3.5 is the one you can self-host; BGE-reranker-v2-m3 is the free default. Voyage reports rerank-3 beating Cohere Rerank v4.0 Pro by 2.72% NDCG@10 across 95 datasets and by over 13% on long documents — but that is Voyage’s benchmark of Voyage’s model, which is reason for caution, not dismissal. Facts verified October 3, 2026.
Side by side
| Voyage rerank-3 | Cohere Rerank 4 | Jina Reranker v3.5 | BGE-reranker-v2-m3 | |
|---|---|---|---|---|
| Released | Sep 30, 2026 | 2026 (Pro + Fast) | Jul 27, 2026 | Open-weight incumbent |
| Price | $0.05 / MTok (lite: $0.02) | $0.0025 / search (Pro) $0.002 (Fast) | Token packages | Free (self-host) |
| Free tier | 200M tokens per account | — | Starter tokens; 10M per new key | Unlimited |
| Context | 32K | 32K | 131K | ~8K typical |
| Parameters | Not disclosed | Not disclosed | 0.6B | ~568M |
| Architecture | Cross-encoder | Joint query+candidate | Listwise | Cross-encoder |
| Self-host | ❌ API only | ❌ API only | ✅ HF + air-gapped | ✅ Apache 2.0 |
| Instruction-following | ✅ | — | — | ❌ |
| Multilingual | ✅ (31 languages tested) | ✅ 100+ languages | ✅ | ✅ |
| Billing unit | All query + document tokens | Query + up to 100 docs | Tokens | Your GPU |
Watch the billing units — they are not comparable at a glance. Cohere’s $0.0025 “search” covers a query plus up to 100 documents regardless of length. Voyage’s $0.05 per million tokens multiplies by every token in query and candidates. For 100 short snippets of 200 tokens each (~20K tokens), Voyage rerank-3 costs about $0.001 — cheaper than Cohere Pro. For 100 long documents at 8K tokens each (~800K tokens), Voyage costs about $0.04, roughly 16× Cohere’s flat $0.0025. Long candidates favour per-search pricing; short snippets favour per-token.
What Voyage actually measured
Voyage’s methodology is unusually transparent, which makes the claims assessable:
- 95 datasets across 9 domains: technical docs, code, law, finance, web reviews, multilingual (51 datasets, 31 languages), long documents, medical, conversations.
- Metric: NDCG@10 after reranking the top 100 candidates from a first-stage retriever.
- Four first-stage retrievers tested: BM25, OpenAI text-embedding-3-large, voyage-3-large, voyage-4-large.
- Baselines: rerank-2.5, Cohere Rerank v4.0 Pro and Fast, Qwen3-Reranker-8B, Jina Reranker v3.5, mxbai-rerank-large-v2, NVIDIA Llama-Nemotron Rerank 1B v2.
Headline results: rerank-3 beats Cohere Rerank v4.0 Pro by 2.72% and Qwen3-Reranker-8B by 3.02%; rerank-3-lite beats them by 2.23% and 2.53%. On LongEmbed (NarrativeQA, SummScreenFD, QMSum), rerank-3 beats rerank-2.5 by 3.35% and Cohere Rerank v4.0 Pro by 13.84%. On code retrieval, rerank-3 improves on rerank-2.5 by 2.17% atop voyage-3-large. Voyage says rerank-3 was top on every first-stage method tested, and that rerank-3-lite matches rerank-2.5 quality at 40% of the price.
The caveats that matter. This is a vendor-run benchmark where three of the four first-stage retrievers are either Voyage’s own or chosen by Voyage, and one of them (voyage-4-large) is the model the headline chart uses. Qwen3-Reranker-8B’s cost was assumed at $0.10/MTok rather than measured. A 2.72% NDCG@10 margin is real but modest — it is not the gap between working and broken retrieval. The 13.84% long-document margin is the genuinely large claim and the one worth testing yourself if long candidates are your workload.
The migration detail nobody mentions
Voyage made rerank-3 a drop-in upgrade: same API, same 32K context, same instruction-following, same price, and — critically — relevance scores calibrated to match the Rerank 2.5 score distributions. If you tuned a score threshold like “drop anything below 0.4” on rerank-2.5, that threshold still works.
This is rarer than it should be and worth demanding from any retrieval vendor. Most reranker upgrades silently shift the score distribution, which means your cutoffs are now wrong and your recall quietly changes without a single error in your logs. Score recalibration is the difference between a one-line model swap and a week of re-tuning.
Which to pick
- Best quality on long documents or code retrieval: Voyage rerank-3. The long-document margin is the biggest claim in the space, and coding-agent queries are explicitly what it was retrained for.
- High volume, short snippets, want flat predictable cost: Cohere Rerank 4 Fast at $0.002 per search. Per-search billing caps your exposure when candidate length varies.
- Broadest language coverage as a managed service: Cohere Rerank 4 — 100+ languages, and the most battle-tested enterprise support story.
- Must self-host, air-gap, or keep data resident: Jina Reranker v3.5. 0.6B parameters, a 131K context window, listwise scoring for lower latency, and weights on Hugging Face. Check the commercial licence before production use.
- Zero budget, needs to be free forever: BGE-reranker-v2-m3, Apache 2.0, runs on a single consumer GPU. Still the most deployed open reranker in production RAG. Expect to fine-tune for out-of-domain corpora.
- Open weights but want more headroom: Qwen3-Reranker at 0.6B / 4B / 8B, strong multilingual accuracy and 32K context on the 8B.
Latency-cheap starting point: rerank-3-lite at $0.02/MTok with 200 million free tokens is a genuinely free way to find out whether reranking helps your corpus before you spend anything.
Related: best reranker models 2026 ranked, how to add a reranker to your RAG pipeline, best embedding models 2026.
Last verified: October 3, 2026. Prices from vendor pricing pages; benchmark figures are Voyage AI’s published results.