Founder Insight

How to Automatically Verify AI Citations and Catch Hallucinations

Ofer Mendelevitch, Author, Independent AI Advisor at O'Reilly — Hands-On RAG for Production

Listen on TL;Listen Prefer to listen? Hear this article read aloud.

Every report a language model writes now ends with a tidy appendix of references. ChatGPT does it, Claude does it, every research tool does it. The references look authoritative. The problem is that when you actually open the source and read the paragraph the model pointed to, it often doesn’t say what the model claimed it said. It’s partially right, or the number is off, or the source is real but the sentence is an extrapolation. Checking it by hand is slow, and it’s the exact work you built the AI to avoid.

That’s the problem Angelina brought to Ofer Mendelevitch on the show, and instead of talking abstractly, they designed a system for it live. Ofer is the author of O’Reilly’s Hands-On RAG for Production and an independent AI advisor who spent years leading developer relations at Vectara, an enterprise RAG platform. He’s been building with language models since 2019, and he’s spent that time on exactly this class of problem: how do you make a retrieval system tell you when a claim isn’t actually supported?

The short answer he kept returning to: plain RAG isn’t enough for this. You need an agent on top of it.

Why a single RAG query can’t verify a document

The intuitive setup is to treat verification as one big retrieval question — hand the whole document to a RAG pipeline and ask it to flag everything wrong. Ofer’s caution is that this collapses under its own weight. “If the number of things is more than three and the document is pretty long, this probably doesn’t work in its one-shot form,” he said. A basic RAG query is “a strict matching of a query to a set of documents,” and a request like fix all the problems in this report doesn’t match any specific chunk of source material.

There’s a sharper failure hiding inside it, one Angelina named before he did: semantic search is bad at numbers. If your source says 90% of POCs fail in production and the generated report says 85%, that’s a mismatch a human would catch instantly. But a vector search can rate the two sentences as a match because almost every other word is identical. “If you look for 85 and 95 differences, they might not show up in semantic search,” Ofer said. “It’s not really a thing.” Verification that leans on similarity alone will wave through the errors that matter most.

The agentic verification loop

The design that works treats the document not as one query but as a list of claims, each checked on its own. This is where agents come in. In Ofer’s framing, an agent is a language model that “reasons about your query, plans what to do, and then executes it by calling one or more tools” — and one of those tools is the RAG retrieval pipeline you already have.

Here’s the shape of it for citation checking:

  • Decompose first. The agent reads the report you want to validate and breaks it into individual claims. “Maybe you have 10 different parts of the document that you want to verify. The LLM should find all these parts on its own — you don’t have to tell it what it is.”
  • Check each claim as its own retrieval. “It will probably call the tool 10 times. Each time it will ask, is this statement correct?” Every claim becomes a targeted query against the source material instead of one vague pass over everything.
  • Pull the evidence, not just a verdict. Because you marked each chunk’s location at ingestion time — “which page and which part of the page” — the system can hand back the exact paragraphs that support a claim, so your spot-check reads the evidence instead of hunting through the original.
  • Reassemble and flag. “It’ll take all this information back and review it together and give you a formal answer” — citations confirmed where the evidence holds, the mismatches highlighted, sometimes with a suggested fix if the correct value exists in your sources.

Ofer’s word for the mature version of this is harness. “Harness is like a better agent — it’s the same thing, just a better implementation.” A good harness “is really good at identifying the pieces of your question, breaking it down to different subtasks.” If you’ve watched Claude Code or Codex work, you’ve seen it: “they even tell you, here’s my five tasks I’m gonna do, and then do them one by one.”

Why this beats a single LLM call

Angelina’s honest description of the naive version is familiar to anyone who’s tried it: hand a complex claim and a source to a language model and ask is this true? “It can only verify half of a sentence,” she said. “Sometimes it’s partial, sometimes it’s mixed, sometimes it’s an extrapolation that you kind of think, maybe I can say that, but when you apply human judgment you have a different read.”

The agentic loop is what dissolves that. Instead of one model straining to hold a whole nuanced claim, plus its numbers, plus its sources in a single pass, the harness splits the claim into pieces small enough to verify cleanly and runs each one to ground. The numeric checks can be routed to logic that actually compares values rather than vibes. The nuanced sub-claims get their own retrieval. You trade one overloaded prompt for many precise ones — which is the whole reason agents exist.

The reframe worth keeping: citation-checking isn’t a search problem, it’s an orchestration problem. Retrieval is a tool the agent calls, not the thing doing the reasoning.

FAQ

Why do AI-generated citations often fail to support the claim?

Language models generate fluent text and attach plausible-looking references, but they don’t rigorously check that each source contains the specific claim. The result is citations that are partially correct, cite real sources for extrapolated points, or misstate numbers — errors a human catches on inspection but the model doesn’t flag on its own.

Can a RAG system verify whether an AI report is accurate?

Basic RAG struggles with it. A one-shot query like “find everything wrong in this document” doesn’t match any specific source chunk, and semantic search misses exact-number mismatches. Verification works better when an agent decomposes the report into individual claims and checks each one as a separate retrieval against the source material.

How do you catch wrong numbers that semantic search misses?

Semantic search can rate “85%” and “95%” as a match because the surrounding words are identical. To catch it, the claim has to be isolated and the number compared directly rather than by similarity. An agentic pipeline separates numeric claims into their own verification step instead of relying on vector similarity to flag the difference.

What is agentic RAG and why does it help with verification?

Agentic RAG puts a reasoning language model in a loop on top of retrieval. It plans, breaks a request into subtasks, and calls the RAG pipeline as a tool — potentially many times. For verification, it splits a document into individual claims and checks each against the sources, instead of trying to validate everything in one pass.

How do you return evidence for each claim instead of just a yes/no?

At ingestion, mark each chunk with its location — which page and which part of the page. Then the verification agent can retrieve and return the exact supporting paragraphs alongside its verdict. You review the evidence directly for a spot-check rather than reopening the source and hunting for the passage yourself.

How many source documents can this scale to?

The design is the same whether you have 100 documents or a million, but the engineering changes. At small scale you re-ingest freely; at production scale you need incremental refresh, careful chunking, and instrumentation. The agent’s per-claim retrieval loop still applies — what gets harder is keeping the underlying index accurate and current.

What is a harness in agentic systems?

A harness is a more robust implementation of an agent loop — the same idea, executed well. It reliably breaks a complex request into subtasks, runs each, and reassembles the results. Coding tools like Claude Code and Codex are visible examples: they announce a task list and work through it step by step, which is the pattern useful for claim-by-claim verification.

Should you build a fact-checking pipeline or use one LLM call?

A single LLM call tends to verify only part of a complex claim and can wave through extrapolations. Splitting the work — an agent that decomposes claims, runs targeted retrieval for each, and separates numeric checks — produces more reliable results. Build the loop when accuracy matters; a lone prompt is fine only for short, simple claims.

Does long context remove the need for this?

No. Even with large context windows, dumping a whole document into one prompt reproduces the one-shot problem: the model reasons over everything at once and misses precise mismatches. The value of the agentic approach is decomposition — checking claims individually — which is an orchestration benefit that a bigger context window doesn’t provide on its own.

Watch the full conversation

Hear Ofer Mendelevitch share the full story on Heroes Behind AI.

Watch on YouTube

More from Ofer Mendelevitch

Founder Archetype

Read Ofer Mendelevitch's archetype profile

· Classical: ·

Related Insights