Understand RAG evaluation.

What a RAG is, why it fails, and how a good evaluation set catches those failures — in plain language. Same three verbs as the rest of our work: generate, validate, score.

RAG in one paragraph

A RAG (Retrieval-Augmented Generation) app answers a question by first retrieving relevant passages from a document corpus, then asking an LLM to generate an answer grounded in them. It's how chatbots answer over private docs. The promise: answers backed by your sources. The risk: when retrieval misses or the corpus has no answer, a weak RAG invents one.

Essential vocabulary

TermIn plain words
corpusThe set of documents the RAG answers from.
retrievalFinding the passages most relevant to a question.
groundingAn answer is grounded when it's actually supported by a retrieved passage.
hallucinationA confident answer that isn't supported by the corpus — the failure that matters most.
eval itemOne test: a question + its golden answer + its gold source + a difficulty label.
golden answerThe correct answer, used to score the RAG's answer.
typologyA family of test with its own difficulty (e.g. no_answer, distractor).
abstentionThe RAG correctly saying "I don't know" when the corpus can't answer.

1 · Generate

From the corpus we build labelled eval items across typologies. The easy ones (answerable) have their answer sitting in a single passage. The hard ones are the point:

  • no_answer — a plausible question the corpus cannot answer. A good RAG must refuse; a weak one hallucinates.
  • distractor — a passage that looks relevant but is wrong competes with the right one.
  • multi_hop — the answer needs two passages combined.

Questions can be written by a rule-based generator (deterministic, no LLM) or by an LLM (natural phrasing) — the same labelled shape either way.

2 · Validate

An eval set is only worth running if its answers are right — a wrong golden answer is worse than no test at all. So a grounding gate drops any item whose golden answer isn't actually supported by its gold source, and removes duplicates. Quality of the test is the moat.

3 · Score

We send every question to your RAG and score the reply on three things:

  • Retrieval — did it fetch the gold source?
  • Answer correctness — does the answer match the golden answer?
  • Refusal — on no_answer, did it correctly abstain instead of inventing?

The result is a per-typology scorecard. The headline is the gap: a RAG can pass answerable perfectly and still fail every no_answer trap.

1.00
answerable · pass
0.60
distractor · pass
0.00
no_answer · abstention

In this example the RAG answered 100% of the questions it should have refused. That single number — the hallucination rate on out-of-scope questions — is often the most important thing to know before shipping a RAG, and the hardest to get without a tailored eval set.

Trust, like our datasets

The eval set and its scorecard ship as a signed bundle (SHA-256 manifest + Ed25519), reproducible from a fixed seed. No real user data is needed to build it, and it's generated with an Apache-2.0 model so the output is yours to use commercially. See how verification works →