Learn / Search quality

How do you evaluate a RAG system?

Updated 3 October 2026 · 2 min read

Short answer

Evaluate a RAG system by building a set of real questions with known good sources and answers, then measuring retrieval (did the right passage come back, and how high) and generation (was the answer correct, supported by the evidence, and refused when it should be).

With troveGEN

troveGEN ships evaluation inside the product, so you measure retrieval and answers on your own documents instead of trusting someone else's benchmark.

See what troveGEN provides ↓

Measure retrieval and answers separately

A wrong answer has two possible causes: the right evidence was not retrieved, or it was retrieved and misused. Measuring them separately tells you which half to fix.

Retrieval metrics

  • Hit rate at k: how often a correct source appears in the top k results.
  • Mean reciprocal rank (MRR): rewards putting the right source near the top.
  • nDCG: a graded measure of ranking quality that accounts for position.
  • Context recall: whether the retrieved text contains the facts needed for a complete answer, which matters when a fact spans several passages.

Answer metrics

  • Correctness: does the answer match the reference answer or contain the required facts.
  • Faithfulness (groundedness): is every claim supported by the retrieved passages.
  • Refusal accuracy: does the system decline questions the documents cannot answer, instead of inventing something.

Build a good question set

Use real questions from users or staff, include hard cases such as acronyms and multi-part questions, and include questions that should be refused. A few dozen well-chosen questions beat thousands of synthetic ones. Questions can also be generated from your documents and then reviewed, which gives a useful starting point.

Make it a habit

Run the same set whenever you change chunking, the embedding model, the reranker or a threshold, and compare. A change that helps one kind of question often hurts another, and only a repeatable test shows it.

Key takeaways

  • Evaluate retrieval and generation separately.
  • Use your own documents and real questions, including ones that should be refused.
  • Re-run the same set after every change.

How troveGEN helps with evaluating RAG

troveGEN includes question sets, scored runs and side-by-side comparison with hit rate, MRR, nDCG, context recall, answer correctness, faithfulness and refusal accuracy. It can draft questions from your own documents for you to review. We publish no leaderboard score, because a number on someone else's documents says little about yours.

What troveGEN provides

  • Question sets with expected sources and answers
  • Draft questions generated from your own documents for you to review
  • Hit rate, MRR, nDCG, context recall, answer correctness, faithfulness and refusal accuracy
  • Side-by-side comparison of two configurations
  • A quality gate before any cheaper model is allowed to handle live questions

Run an evaluation free Start free — 500 pages

Frequently asked questions

How many test questions do I need?

Start with 30 to 50 representative ones. Add every failure you meet in production so the set grows with your experience.

Can I use an AI model to judge answers?

Yes, as one signal, ideally combined with reference answers and mechanical checks. Review a sample by hand to make sure the judge agrees with you.

What is a good score?

It depends on your documents and tolerance for errors. Use scores to compare configurations and track change over time rather than to chase a universal number.

How does troveGEN help with evaluating RAG?

troveGEN ships evaluation inside the product, so you measure retrieval and answers on your own documents instead of trusting someone else's benchmark. It provides: Question sets with expected sources and answers; Draft questions generated from your own documents for you to review; Hit rate, MRR, nDCG, context recall, answer correctness, faithfulness and refusal accuracy; Side-by-side comparison of two configurations; A quality gate before any cheaper model is allowed to handle live questions.

Keep reading