Reranking

First-stage retrieval (vector search, keyword search, or hybrid search) is built for speed over millions of documents. It compares a query against precomputed document representations and returns candidates quickly, but its ranking is approximate. A reranker is a second, slower, more accurate model that looks at the query and each candidate together and reorders them by true relevance. Retrieve 50–100 candidates cheaply, rerank them precisely, and send only the best few to the LLM.

Adding a reranker is often the single highest-impact improvement to a RAG pipeline: better context, fewer tokens, less distraction, and fewer hallucinations. The cost is added latency (tens to hundreds of milliseconds) and compute.

TL;DR

Quick Example

Retrieve with hybrid search, then rerank with a cross-encoder (sentence-transformers):

Or with a hosted API:

Core Concepts

Bi-Encoders vs Cross-Encoders

A bi-encoder must compress a whole passage into one vector before seeing the query. A cross-encoder reads both at once, so its attention can connect "refund window for digital purchases" with the specific sentence about digital goods. See transformer architecture.

The Two-Stage Pipeline

Reranker Options

LLM-Based Reranking

An LLM can judge relevance with instructions ("rank these passages by how well they answer the question, considering the product version"). Approaches:

Small language models and fine-tuned rerankers often match large LLMs on this task at a fraction of the cost. Use LLM reranking where domain reasoning matters and latency allows.

Relevance Thresholds

Reranker scores are more meaningful than raw similarity scores, so you can set a minimum relevance and pass fewer chunks when fewer are relevant, or trigger "I couldn't find that" behavior when none pass. Calibrate thresholds on evaluation data, because scores vary by model.

Best Practices

Retrieve Wide, Rerank Narrow

Feed the reranker enough candidates (commonly 50–100) from hybrid search so it has the right passages to choose from, then keep a small, high-quality set for the prompt.

Rerank With the Same Text the LLM Will See

If chunks carry contextual headers (title, section path), include them when reranking. They help the reranker, and they're what the model reads. See chunking.

Budget Latency Explicitly

Measure p95 reranking time at your candidate count and document length. Truncate long passages to the reranker's max length, batch pairs, run on GPUs for self-hosted models, and cache results for repeated queries.

Evaluate With Labeled Data

Compare retrieval with and without reranking using nDCG@k, MRR, and recall@k, plus end-to-end answer quality. Rerankers usually help a lot, but the size of the gain, and the best candidate count, is corpus-specific. See RAG evaluation.

Common Mistakes

Reranking Too Few Candidates

The reranker can't recover documents the first stage didn't return. Retrieve more.

Using Reranker Scores Across Models Interchangeably

Score scales and distributions differ between rerankers, and even between versions. Recalibrate thresholds whenever you switch models.

Skipping Reranking to Save Latency, Then Sending 20 Chunks

Sending many loosely relevant chunks costs more tokens and time in generation than a fast reranker costs, and it lowers answer quality. A small reranker plus fewer chunks is often faster end to end.

FAQ

Why not use a cross-encoder for all retrieval?

Cross-encoders must process each query-document pair at query time; they can't precompute document representations. Scoring millions of documents per query would be far too slow. They're practical only on a shortlist produced by fast first-stage retrieval.

How much does reranking improve RAG?

It varies by corpus and first-stage quality, but improvements in retrieval precision are often substantial, and they translate to better answers with fewer context tokens. Anthropic's contextual retrieval experiments, for example, showed reranking further cutting retrieval failures on top of hybrid search.

Which reranker should I use?

For fast adoption, a hosted rerank API with good multilingual support. For control and cost at scale, an open cross-encoder like the BGE or mixedbread rerankers, self-hosted. Benchmark two or three candidates on your own queries: public leaderboards don't always reflect domain performance.

Can an LLM replace the reranker?

It can act as one, and it's excellent when relevance requires reasoning, but it's slower and more expensive per candidate. A common compromise is a cross-encoder for the main reranking step and an optional LLM check on the final few passages.

Related Topics

References