Reranking
First-stage retrieval (vector search, keyword search, or hybrid search) is built for speed over millions of documents. It compares a query against precomputed document representations and returns candidates quickly, but its ranking is approximate. A reranker is a second, slower, more accurate model that looks at the query and each candidate together and reorders them by true relevance. Retrieve 50–100 candidates cheaply, rerank them precisely, and send only the best few to the LLM.
Adding a reranker is often the single highest-impact improvement to a RAG pipeline: better context, fewer tokens, less distraction, and fewer hallucinations. The cost is added latency (tens to hundreds of milliseconds) and compute.
TL;DR
- Bi-encoders (embedding models) encode query and document separately: fast and indexable, but less precise.
- Cross-encoders (rerankers) encode query and document jointly, so attention across both gives far more accurate relevance.
- Standard pipeline: retrieve top 50–200 → rerank → keep top 3–10 for the prompt.
- Options: hosted rerank APIs (Cohere, Voyage, Jina), open models (BGE reranker, mixedbread, ms-marco cross-encoders), or LLM-based reranking.
- Reranking also enables relevance thresholds, so you can drop weak passages instead of always stuffing k chunks.
- Measure gains with nDCG, MRR, and recall@k on labeled queries, and watch latency budgets.
Quick Example
Retrieve with hybrid search, then rerank with a cross-encoder (sentence-transformers):
Or with a hosted API:
Core Concepts
Bi-Encoders vs Cross-Encoders
A bi-encoder must compress a whole passage into one vector before seeing the query. A cross-encoder reads both at once, so its attention can connect "refund window for digital purchases" with the specific sentence about digital goods. See transformer architecture.
The Two-Stage Pipeline
- Recall is the first stage's job: make sure relevant passages are somewhere in the candidates.
- Precision is the reranker's job: put the best passages first.
- Tune the candidate count: more candidates improve recall but add reranking latency roughly linearly.
Reranker Options
LLM-Based Reranking
An LLM can judge relevance with instructions ("rank these passages by how well they answer the question, considering the product version"). Approaches:
- Pointwise: score each passage independently (0–10). Easy to parallelize.
- Listwise: give a list and ask for an ordering. Captures relative relevance, but is sensitive to order and context limits.
- Pairwise: compare two at a time. Accurate but expensive.
Small language models and fine-tuned rerankers often match large LLMs on this task at a fraction of the cost. Use LLM reranking where domain reasoning matters and latency allows.
Relevance Thresholds
Reranker scores are more meaningful than raw similarity scores, so you can set a minimum relevance and pass fewer chunks when fewer are relevant, or trigger "I couldn't find that" behavior when none pass. Calibrate thresholds on evaluation data, because scores vary by model.
Best Practices
Retrieve Wide, Rerank Narrow
Feed the reranker enough candidates (commonly 50–100) from hybrid search so it has the right passages to choose from, then keep a small, high-quality set for the prompt.
Rerank With the Same Text the LLM Will See
If chunks carry contextual headers (title, section path), include them when reranking. They help the reranker, and they're what the model reads. See chunking.
Budget Latency Explicitly
Measure p95 reranking time at your candidate count and document length. Truncate long passages to the reranker's max length, batch pairs, run on GPUs for self-hosted models, and cache results for repeated queries.
Evaluate With Labeled Data
Compare retrieval with and without reranking using nDCG@k, MRR, and recall@k, plus end-to-end answer quality. Rerankers usually help a lot, but the size of the gain, and the best candidate count, is corpus-specific. See RAG evaluation.
Common Mistakes
Reranking Too Few Candidates
The reranker can't recover documents the first stage didn't return. Retrieve more.
Using Reranker Scores Across Models Interchangeably
Score scales and distributions differ between rerankers, and even between versions. Recalibrate thresholds whenever you switch models.
Skipping Reranking to Save Latency, Then Sending 20 Chunks
Sending many loosely relevant chunks costs more tokens and time in generation than a fast reranker costs, and it lowers answer quality. A small reranker plus fewer chunks is often faster end to end.
FAQ
Why not use a cross-encoder for all retrieval?
Cross-encoders must process each query-document pair at query time; they can't precompute document representations. Scoring millions of documents per query would be far too slow. They're practical only on a shortlist produced by fast first-stage retrieval.
How much does reranking improve RAG?
It varies by corpus and first-stage quality, but improvements in retrieval precision are often substantial, and they translate to better answers with fewer context tokens. Anthropic's contextual retrieval experiments, for example, showed reranking further cutting retrieval failures on top of hybrid search.
Which reranker should I use?
For fast adoption, a hosted rerank API with good multilingual support. For control and cost at scale, an open cross-encoder like the BGE or mixedbread rerankers, self-hosted. Benchmark two or three candidates on your own queries: public leaderboards don't always reflect domain performance.
Can an LLM replace the reranker?
It can act as one, and it's excellent when relevance requires reasoning, but it's slower and more expensive per candidate. A common compromise is a cross-encoder for the main reranking step and an optional LLM check on the final few passages.
Related Topics
- RAG — Retrieval-augmented generation overview
- Hybrid Search — The first stage that feeds the reranker
- RAG Chunking — The units being reranked
- RAG Evaluation — Measuring reranking gains
- Embeddings — Bi-encoder retrieval
- Agentic RAG — Iterative retrieval driven by the model