RAG & LLM evaluation

Find out where your RAG breaks.

Your RAG demos beautifully — then confidently answers a question your documents can't support. Plexoria builds hard-case evaluation sets from your corpus and scores your RAG on the failure modes that matter.

Passing a demo isn't passing a test.

A RAG that nails your happy-path questions can still hallucinate on the ones your corpus can't answer, latch onto a plausible-but-wrong passage, or miss a fact spread across two documents. Generic public benchmarks don't cover your corpus — so they can't catch your failures.

Hallucination

The costliest failure: answering a question the corpus doesn't support instead of saying "I don't know".

Distractors

Retrieval grabs a lexically-similar but wrong passage and answers from it.

Multi-hop

The answer needs two passages combined; the RAG stops at one.

A labelled eval set — and a scorecard of your RAG.

Each item is a question with a golden answer and its gold source document, tagged with a difficulty typology. We point the set at your RAG's endpoint and score answer correctness, retrieval, and — the flagship — whether it correctly refuses the unanswerable.

TypologyKindWhat it tests
answerablebaselineThe answer is in one passage — does the RAG find it?
multi_hopovertNeeds two passages combined.
distractorhardA similar-looking wrong passage competes with the right one.
no_answerhardThe corpus can't answer it — the RAG must refuse, not invent.
1.00
answerable · pass
0.60
distractor · pass
0.00
no_answer · abstention

Example scorecard of a naive baseline RAG. The gap between answerable and the hard typologies is the point: it's exactly what a demo hides and a spec-sheet can't show.

Generate · validate · score.

The same rigor as our datasets, aimed at RAG. Reproducible, and signed end to end.

01 Generate

From your corpus we generate labelled items across the typologies — including the hard cases most sets skip.

02 Validate

A grounding gate drops any item whose golden answer isn't supported by its gold source. A wrong golden is worse than useless.

03 Score

We run the set against your RAG endpoint and report a per-typology scorecard: retrieval, answer correctness, and refusal.

04 Sign

Eval set + scorecard ship as a signed bundle (SHA-256 manifest + Ed25519) — reproducible and verifiable.

Tailored to your corpus — you own it.

An eval set is only useful when it's grounded in your documents, so we generate it for your corpus and you own both. No generic file that fits nobody. Governance-friendly: generated with an Apache-2.0 model (commercial output rights), no real user data required, and every bundle is signed. A fit for regulated and sovereign environments.