Understand RAG evaluation.
What a RAG is, why it fails, and how a good evaluation set catches those failures — in plain language. Same three verbs as the rest of our work: generate, validate, score.
RAG in one paragraph
A RAG (Retrieval-Augmented Generation) app answers a question by first retrieving relevant passages from a document corpus, then asking an LLM to generate an answer grounded in them. It's how chatbots answer over private docs. The promise: answers backed by your sources. The risk: when retrieval misses or the corpus has no answer, a weak RAG invents one.
Essential vocabulary
| Term | In plain words |
|---|---|
corpus | The set of documents the RAG answers from. |
retrieval | Finding the passages most relevant to a question. |
grounding | An answer is grounded when it's actually supported by a retrieved passage. |
hallucination | A confident answer that isn't supported by the corpus — the failure that matters most. |
eval item | One test: a question + its golden answer + its gold source + a difficulty label. |
golden answer | The correct answer, used to score the RAG's answer. |
typology | A family of test with its own difficulty (e.g. no_answer, distractor). |
abstention | The RAG correctly saying "I don't know" when the corpus can't answer. |
1 · Generate
From the corpus we build labelled eval items across typologies. The easy ones (answerable) have their answer sitting in a single passage. The hard ones are the point:
no_answer— a plausible question the corpus cannot answer. A good RAG must refuse; a weak one hallucinates.distractor— a passage that looks relevant but is wrong competes with the right one.multi_hop— the answer needs two passages combined.
Questions can be written by a rule-based generator (deterministic, no LLM) or by an LLM (natural phrasing) — the same labelled shape either way.
2 · Validate
An eval set is only worth running if its answers are right — a wrong golden answer is worse than no test at all. So a grounding gate drops any item whose golden answer isn't actually supported by its gold source, and removes duplicates. Quality of the test is the moat.
3 · Score
We send every question to your RAG and score the reply on three things:
- Retrieval — did it fetch the gold source?
- Answer correctness — does the answer match the golden answer?
- Refusal — on
no_answer, did it correctly abstain instead of inventing?
The result is a per-typology scorecard. The headline is the gap: a RAG can pass answerable perfectly and still fail every no_answer trap.
In this example the RAG answered 100% of the questions it should have refused. That single number — the hallucination rate on out-of-scope questions — is often the most important thing to know before shipping a RAG, and the hardest to get without a tailored eval set.
Trust, like our datasets
The eval set and its scorecard ship as a signed bundle (SHA-256 manifest + Ed25519), reproducible from a fixed seed. No real user data is needed to build it, and it's generated with an Apache-2.0 model so the output is yours to use commercially. See how verification works →