RAG Evaluation
A RAG system has many moving parts (parsing, chunking, embeddings, hybrid search, reranking, prompting, and the model), and a change to any of them can quietly improve some answers while breaking others. Evaluation makes those trade-offs visible. Without it, teams tune by anecdote ("this answer looks better now") and ship regressions.
Effective RAG evaluation separates retrieval quality (did we find the right passages?) from generation quality (did the model answer correctly and faithfully from them?), uses a golden dataset of realistic questions, combines deterministic metrics with calibrated LLM-as-judge scoring, and keeps running in CI and production.
TL;DR
- Evaluate retrieval and generation separately. Bad answers from good context and good answers despite bad context need different fixes.
- Build a golden dataset: real questions, expected answers, and the passages that support them, including unanswerable questions.
- Retrieval metrics: recall@k, precision@k, MRR, nDCG, and context precision and recall.
- Generation metrics: faithfulness/groundedness, answer relevance, correctness, citation accuracy, and abstention quality.
- LLM-as-judge scales evaluation. Validate it against human labels, and use rubrics and structured outputs.
- Run evals in CI on every pipeline change, and sample production traffic for ongoing monitoring.
Quick Example
A minimal evaluation loop: retrieval recall plus a judge for faithfulness and correctness.
Core Concepts
The Golden Dataset
A good evaluation set is the most valuable asset in a RAG project:
- Real questions: from support tickets, search logs, user interviews, and domain experts, not only ones an LLM invented.
- Reference answers and supporting passage IDs, so retrieval and generation can be scored separately.
- Coverage: common and rare questions, different document types, multi-hop questions, time-sensitive ones, and unanswerable ones that test abstention.
- Size: 50–200 carefully labeled examples catch most regressions; grow the set as failures are found.
Synthetic question generation (an LLM writing questions from chunks) helps bootstrap coverage, but it skews toward easy, lexically similar questions. Mix it with real data.
Retrieval Metrics
Generation Metrics
LLM-as-Judge
Model graders scale evaluation to thousands of examples. Make them reliable:
- Use specific rubrics with binary or small-scale criteria ("every claim supported: yes/no") rather than vague 1–10 scores.
- Give the judge the context and reference, and ask for reasoning before the verdict, output as structured JSON.
- Validate the judge against a human-labeled sample, and track agreement.
- Watch for biases: preference for longer answers, for their own model family's style, and position bias in pairwise comparisons.
See LLM evaluation.
Tooling
Frameworks and platforms that implement these metrics and workflows include Ragas, DeepEval, TruLens, Promptfoo, Arize Phoenix, LangSmith, Braintrust, and cloud evaluation services. They add dataset management, experiment comparison, and tracing. The metrics matter more than the tool; start with simple scripts if that gets evaluation running today.
Evaluation in the Development Loop
- Baseline: measure the current pipeline on the golden set.
- Change one thing: chunk size, embedding model, top-k, reranker, prompt, or model.
- Compare: retrieval and generation metrics side by side, plus cost and latency.
- Inspect failures: read the examples that got worse, not just the averages.
- Gate in CI: fail builds when key metrics regress beyond a threshold.
- Monitor production: sample real traffic for judge scoring, track user feedback (thumbs, corrections, escalations), and add new failures to the golden set.
See LLMOps for the operational side.
Best Practices
Diagnose by Stage
Low recall means fix retrieval (chunking, hybrid search, query rewriting, metadata filters). High recall but unfaithful answers means fix prompting (grounding instructions, citations, allowing abstention) or the model. Stage-level metrics tell you where to look.
Include Unanswerable Questions
A system that always answers looks great until users ask something outside the corpus. Measure how often it correctly abstains, and how often it wrongly refuses.
Track Cost and Latency Alongside Quality
A change that adds 2% correctness but doubles latency or token cost may not be worth it. Report tokens per answer, p95 latency, and quality together.
Version Everything
Record dataset version, pipeline configuration, prompts, model versions, and judge version for every run, so results are reproducible and comparable over time.
Common Mistakes
Judging Only the Final Answer
End-to-end scores hide why quality changed. Always log the retrieved chunks and score retrieval independently.
Evaluating on Synthetic Questions Alone
LLM-generated questions often reuse the source passage's wording, making retrieval look far better than it is for real users' phrasing. Include real queries.
Trusting an Unvalidated Judge
A judge that disagrees with humans 30% of the time produces confident but misleading metrics. Spot-check judge verdicts regularly, and recalibrate prompts when you change judge models.
FAQ
What are the most important RAG metrics?
Retrieval recall@k (is the needed information in the context?) and faithfulness (is the answer supported by that context?), plus correctness against reference answers. Add answer relevance, citation accuracy, and abstention quality as your system matures.
How many examples do I need in an evaluation set?
Start with 50–100 high-quality, diverse examples. That's enough to catch major regressions. Grow toward several hundred as you discover failure modes, and keep subsets per category (product area, question type) to spot localized regressions.
Is LLM-as-judge reliable?
Reasonably, when designed well: specific rubrics, reference answers and context provided, binary criteria, reasoning before verdicts, and validation against human labels. It's not a replacement for periodic human review, especially in high-stakes domains.
How do I evaluate RAG in production without labels?
Use reference-free metrics (faithfulness to retrieved context, answer relevance, context relevance) judged by an LLM on sampled traffic, plus behavioral signals: user feedback, follow-up rephrasing, escalations to humans, and click-through on citations. Feed discovered failures back into the golden set.
Related Topics
- RAG — Retrieval-augmented generation overview
- LLM Evaluation — Evaluating LLM systems in general
- LLMOps — Running evaluation and monitoring in production
- LLM Hallucinations — What faithfulness metrics catch
- Reranking — A common lever evaluation measures
- RAG Chunking — Another lever to tune with evals