RAG & LLM evaluation
Find out where your RAG breaks.
Your RAG demos beautifully — then confidently answers a question your documents can't support. Plexoria builds hard-case evaluation sets from your corpus and scores your RAG on the failure modes that matter.
Passing a demo isn't passing a test.
A RAG that nails your happy-path questions can still hallucinate on the ones your corpus can't answer, latch onto a plausible-but-wrong passage, or miss a fact spread across two documents. Generic public benchmarks don't cover your corpus — so they can't catch your failures.
Hallucination
The costliest failure: answering a question the corpus doesn't support instead of saying "I don't know".
Distractors
Retrieval grabs a lexically-similar but wrong passage and answers from it.
Multi-hop
The answer needs two passages combined; the RAG stops at one.
A labelled eval set — and a scorecard of your RAG.
Each item is a question with a golden answer and its gold source document, tagged with a difficulty typology. We point the set at your RAG's endpoint and score answer correctness, retrieval, and — the flagship — whether it correctly refuses the unanswerable.
| Typology | Kind | What it tests |
|---|---|---|
answerable | baseline | The answer is in one passage — does the RAG find it? |
multi_hop | overt | Needs two passages combined. |
distractor | hard | A similar-looking wrong passage competes with the right one. |
no_answer | hard | The corpus can't answer it — the RAG must refuse, not invent. |
Example scorecard of a naive baseline RAG. The gap between answerable and the hard typologies is the point: it's exactly what a demo hides and a spec-sheet can't show.
Generate · validate · score.
The same rigor as our datasets, aimed at RAG. Reproducible, and signed end to end.
01 Generate
From your corpus we generate labelled items across the typologies — including the hard cases most sets skip.
02 Validate
A grounding gate drops any item whose golden answer isn't supported by its gold source. A wrong golden is worse than useless.
03 Score
We run the set against your RAG endpoint and report a per-typology scorecard: retrieval, answer correctness, and refusal.
04 Sign
Eval set + scorecard ship as a signed bundle (SHA-256 manifest + Ed25519) — reproducible and verifiable.
Tailored to your corpus — you own it.
An eval set is only useful when it's grounded in your documents, so we generate it for your corpus and you own both. No generic file that fits nobody. Governance-friendly: generated with an Apache-2.0 model (commercial output rights), no real user data required, and every bundle is signed. A fit for regulated and sovereign environments.