How to Evaluate a RAG System Without a Golden Answer Set
Ofer Mendelevitch, Author, Independent AI Advisor at O'Reilly — Hands-On RAG for Production
Ask most teams how they evaluate their RAG system and you’ll hear about a golden dataset — a curated set of queries, each paired with the correct answer or the right source chunks, that you score against. It’s the textbook approach. It’s also the one that quietly falls apart in production, because building and maintaining that dataset is often harder than building the RAG system itself.
Ofer Mendelevitch has watched this play out repeatedly. He’s the author of O’Reilly’s Hands-On RAG for Production and an independent AI advisor who led developer relations at Vectara, an enterprise RAG platform. The evaluation chapter of his book spends less time on which metric to compute and more on the problem underneath: where does the ground truth come from when you have a million documents that change every week? His answer points to a body of research — reference-free evaluation — that sidesteps the golden set entirely.
The real problem isn’t the metric, it’s the data
When you evaluate RAG, there are two things to check. First, retrieval: “Did I actually get the right chunks of information?” Second, generation: was the final answer correct and did it actually use what was retrieved. The metrics for the retrieval half are well understood — precision, recall, the classics.
“But the real problem is not the metrics,” Ofer said. “It’s how do you collect the data for this.”
Picture what a golden set demands. For a thousand evaluation queries, someone has to specify, by hand, the exact chunks that should come back for each one — not the document, the chunks. “It’s really hard to sit down and write that up,” and it’s the same tedious source-hunting that made you want to automate the system in the first place. Then production makes it worse: “imagine that you also have a million documents, not a hundred, and then the documents always change.” The moment your corpus updates, the correct chunk for a query can be different, and your painstakingly labeled dataset is stale. Some teams throw people at it — “two or three people for a few weeks” writing curated answers — and it’s still imperfect, because there are five valid ways to phrase the same response.
That’s the trap reference-free evaluation is built to escape.
UMBRELA: scoring chunks without an answer key
The work Ofer points to comes out of Jimmy Lin’s lab at the University of Waterloo, and it’s covered in the book. “This idea of reference-free evaluation of RAG” gives you metrics that don’t require a golden answer or a golden chunk, yet still produce a trustworthy quality score.
The retrieval-side metric he walks through is called UMBRELA. The mechanic is simple to state:
- Take the chunks your system retrieved for a query — say the top 10.
- For each chunk, make an LLM call that scores its relevance to the query on a 0-to-3 scale.
- 0 — the chunk has nothing to do with the query.
- 1 — it’s tangentially related.
- 2 — quite relevant, but doesn’t fully answer.
- 3 — a perfect, complete answer.
- Aggregate those scores into a picture of how good your retrieval was — no labeled ground truth required.
If that sounds like plain LLM-as-a-judge, that’s the honest catch, and Ofer says so directly: “You can always judge things with a prompt.” Scoring text with a language model is not novel. So what makes UMBRELA worth using instead of a prompt you improvised on a Tuesday?
Why the validation is the whole point
The value isn’t the 0-to-3 rubric. It’s the evidence behind it. “The team has done research and asked human beings to also rank these chunks, and they’ve shown that the correlation of that particular prompting technique highly correlates with the human answers. That’s the power of it.”
Read that carefully, because it’s the part teams skip. Anyone can prompt an LLM to score relevance. What you can’t easily do yourself is prove that your scoring prompt agrees with human judgment — which is the only reason to trust the number at all. Without that validation, “you could have biases of LLMs and all kinds of different things” contaminating your eval, and you’d never know. The Waterloo work did the human-correlation studies across multiple datasets so you don’t have to. You’re not borrowing a rubric; you’re borrowing the proof that the rubric tracks reality.
The practical takeaway for anyone building RAG: stop treating “we don’t have a golden set” as a reason not to evaluate. Reference-free methods let you measure retrieval quality continuously, even as your million documents churn — and the ones grounded in human-correlation research are the ones you can actually trust. The generation side has analogous techniques, also validated against human agreement rather than a fixed answer key.
FAQ
What is reference-free RAG evaluation?
It’s a way to score a RAG system’s quality without a golden answer set or labeled correct chunks. Instead of comparing outputs to a hand-curated ground truth, it uses validated LLM-based scoring to judge whether retrieved chunks are relevant and whether generated answers are good, so you can evaluate even when maintaining a labeled dataset is impractical.
Why is building a golden dataset for RAG so hard?
You have to specify, by hand, the exact chunks that should be retrieved for each of hundreds or thousands of queries — the same tedious source-hunting RAG was meant to automate. In production, with a million documents that constantly change, the correct chunks shift, so a labeled dataset goes stale almost as fast as you build it.
What is UMBRELA?
UMBRELA is a reference-free metric for the retrieval side of RAG from Jimmy Lin’s lab at the University of Waterloo. For each retrieved chunk, it uses an LLM call to score relevance to the query from 0 (unrelated) to 3 (a perfect answer). Aggregating those scores estimates retrieval quality without any labeled ground truth.
How does the 0-to-3 chunk scoring work?
Each retrieved chunk gets an LLM-assigned score: 0 means no relation to the query, 1 means loosely related, 2 means quite relevant but incomplete, and 3 means a complete, perfect answer. Run it over your top retrieved chunks per query and aggregate to gauge how well retrieval is performing — no golden chunks needed.
Isn’t this just LLM-as-a-judge?
Mechanically, yes — scoring text with an LLM prompt is not new. The difference is validation. The Waterloo researchers had humans rank the same chunks across multiple datasets and showed their prompting technique highly correlates with human judgment. That human-correlation evidence is what makes the automated scores trustworthy rather than just plausible.
Why does human correlation matter for LLM-based evaluation?
Because an LLM judge can carry hidden biases you won’t detect on your own. If your scoring prompt doesn’t actually track what humans consider relevant, your eval is misleading. Validating the technique against human rankings proves the automated score reflects real quality — turning “the model said so” into a defensible measurement.
How do you evaluate the generation part of a RAG system?
You check whether the final answer is correct and whether it properly uses the retrieved results, since the model can ignore them or add its own material. Reference-free techniques exist here too, also validated against human agreement, so you can score answer quality without maintaining a fixed golden-response set for every query.
Can you evaluate RAG when your document set changes constantly?
That’s exactly where reference-free evaluation shines. Golden datasets break when documents change because the correct chunks move. Because reference-free scoring judges relevance at evaluation time rather than against fixed labels, you can keep measuring retrieval and generation quality continuously as the underlying corpus updates.
Watch the full conversation
Hear Ofer Mendelevitch share the full story on Heroes Behind AI.
Watch on YouTubeMore from Ofer Mendelevitch
Founder Archetype
Read Ofer Mendelevitch's archetype profile
· Classical: ·
Related Insights
How to Test LLM Applications: The Behavioral Testing Framework
Ali Parandeh, Head of Engineering / Author at Building Generative AI Services with FastAPI (O'Reilly)
Why 80% of RAG Pipelines Fail in Production
Gil Feig, CTO at Merge
AI Agents Are Distributed State Machines — What That Means for How You Build Them
Nicole Königstein, Founder at AgensFlow