Why RAG Systems Break in Production (When Demos Worked Fine)
Ofer Mendelevitch, Author, Independent AI Advisor at O'Reilly — Hands-On RAG for Production
The RAG demo everyone remembers is the same one. You drop a PDF into a vector store, wire it to a language model, and suddenly you can ask a 200-page document questions and get grounded answers. “Look mom, my hands are free — I can talk to my PDF,” as Ofer Mendelevitch puts it. It feels magical. Then you point the same pipeline at your company’s actual document set and it starts failing in ways the demo never hinted at.
Ofer wrote O’Reilly’s Hands-On RAG for Production because that gap is where almost all the real work lives. He’s an independent AI advisor who led developer relations at Vectara, an enterprise RAG platform, and has been building with language models since 2019 — which means he’s watched a lot of teams ship a magical prototype and then discover that “a lot of the stuff was just prototypes, just initial demos. It is very helpful to build systems with this, but when you scale them to production, a lot of other things start to break.”
Here’s what actually breaks, and why the “RAG is dead” headline misreads the problem.
”RAG is dead” is a marketing cycle, not a diagnosis
Before the failure modes, the framing. Every few months a fresh wave of “RAG is dead” posts arrives, usually pinned to long-context models that can swallow a million tokens. Ofer’s read is blunt: “There’s been RAG is dead claims since the last two years, every two months — people mostly trying to market something new.”
His actual argument isn’t that RAG is immortal, it’s that people mistake the acronym for one specific 2023 demo. “If you think about RAG, it’s three components: retrieval, augmented, generation. The R is for retrieval, and retrieval is still extremely important, even with agents.” What dies is the naive one-shot version — retrieve once, generate once, ship. What replaces it is retrieval used as a tool inside an agent loop. The retrieval doesn’t go away. It goes underneath.
That’s the reframe that matters when your CEO asks whether the vector database is still necessary. The answer is that finding the right information at the right moment gets more important as systems get more autonomous, not less.
The failure modes the demo hides
Once you’re past a handful of clean PDFs, the breakages are concrete and repeatable:
- Semantic search misses exact values. The one that surprises people most. “If you look for 85 and 95 differences, they might not show up in semantic search. It’s not really a thing.” Two sentences that differ only in a number read as a match, because everything else is identical. For anything numeric — statistics, prices, dates — pure vector similarity is the wrong tool, which is why hybrid search (adding BM25 or TF-IDF) exists.
- Tables get shredded by chunking. Standard chunking splits text into pieces, and a table doesn’t survive that. “The first chunk will have column names and five rows, and then the second chunk will not have column names — so now the LLM doesn’t know what it means.” The fix is to stop treating tables as text: “you gotta treat the table as a first-class citizen,” store it whole, and hand the full table to the model at query time.
- Video and images lose their meaning in transcription. Ofer’s vivid example is the “red button problem.” A training video where the instructor says, “in order to avoid catastrophe, never press this button” — transcribe the audio alone and “you lose the meaning completely,” because this button only makes sense with the picture. Multimodal content needs the visual understood, not just the words captured.
- Ingestion falls over on real files. “Some companies have documents that are 15,000 pages.” Build a naive pipeline and “page 4,000, it goes out of memory and you have to restart it.” That single detail — the OOM at page 4,000 — “is the difference between a prototype and a production system.”
None of these show up when you’re demoing with two clean PDFs. All of them show up in week two of a real deployment.
Cost and maintenance are where it gets expensive
The part teams underestimate is that a RAG index is not a one-time build. Documents change. “Index all of your Google Drive… Google Drive documents change every day.” At ten documents you just re-ingest. At a million, re-ingesting everything nightly is untenable, so you need an incremental strategy — either hashing documents and only re-processing the ones that changed, or a trigger-based refresh where the source system tells you “these five files have changed, go re-index them.”
On raw cost, Ofer’s point is that the first ingest dominates. “The first ingest is usually the most expensive one” — especially with a million complex PDFs full of tables and images. Refreshes afterward are cheap because you’re only touching the small percentage that moved. The real spend isn’t API calls; it’s the engineering to make ingestion “very scalable” and “handle messy data, lots of data, long data.” His advice is to “build with flexibility initially, and a lot of instrumentation, so you can start scaling your data and see what happens” — because “you always find these edge cases. This type of file with this type of image always fails, and I have to fix it. It’s a continuous journey.”
That’s the honest version of the demo-to-production story. The magic was real. Production is just a different job.
FAQ
Why does a RAG system that worked in a demo fail in production?
Demos use a few clean documents; production means millions of messy ones in mixed formats. Chunking mangles tables, semantic search misses exact numbers, multimodal content loses meaning in transcription, and ingestion pipelines run out of memory on very large files. The retrieval logic is the same — the scale and data quality expose failure modes the demo never touched.
Is RAG dead now that models have long context windows?
No. “RAG is dead” claims recur every couple of months, usually to market something new. The naive retrieve-once-generate-once pattern is fading, but retrieval itself is more important with agents, not less — it becomes a tool the agent calls repeatedly. Long context doesn’t remove the need to find the right information precisely.
Why does semantic search miss numbers like 85% vs 95%?
Vector similarity scores sentences by overall meaning. Two sentences that differ only in a number are almost identical in every other word, so they register as a match. For numeric or exact-token queries, similarity is the wrong signal — you need hybrid search that combines semantic search with keyword methods like BM25 or TF-IDF.
How do you handle tables in a RAG pipeline?
Don’t let standard chunking split them. A table cut into chunks loses its column headers partway through, and the model can’t interpret the rows. Instead, detect the table, store it as a whole unit — a first-class citizen — and retrieve and pass the entire table to the language model at query time so it can reason over the full structure.
What is the “red button problem” in multimodal RAG?
It’s Ofer’s example of why transcription isn’t enough for video. An instructor saying “never press this button” is meaningless without seeing which button — transcribe the audio alone and you lose the point entirely. Multimodal RAG has to understand the visual content and correlate it with the audio, not just capture the spoken words.
How do you keep a RAG index up to date without re-ingesting everything?
Two common approaches. Store a hash of each document and, on a scheduled run, only re-process the ones whose hash changed. Better, when the source supports it, use trigger-based refresh: the system (Google Drive, S3) notifies your pipeline that specific files changed so you re-index just those. Full nightly re-ingestion doesn’t scale to a million documents.
Which part of a RAG system costs the most?
The first ingest, especially for large volumes of complex documents with tables, images, or video. Subsequent refreshes are cheap because they touch only the small share of documents that changed. The bigger hidden cost is engineering — building an ingestion pipeline scalable and resilient enough to handle messy, oversized, edge-case-laden real-world data.
What separates a RAG prototype from a production system?
Resilience to real data. A prototype handles a few clean PDFs; a production system handles 15,000-page files without running out of memory, mixed formats, constant document changes, and endless edge cases. Ofer’s recommendation is to build with flexibility and heavy instrumentation from the start, then scale the data gradually and fix what breaks — it’s a continuous process, not a one-time build.
Watch the full conversation
Hear Ofer Mendelevitch share the full story on Heroes Behind AI.
Watch on YouTubeMore from Ofer Mendelevitch
Founder Archetype
Read Ofer Mendelevitch's archetype profile
· Classical: ·
Related Insights
Do You Actually Need a Knowledge Graph for Your RAG System?
Ofer Mendelevitch, Author, Independent AI Advisor at O'Reilly — Hands-On RAG for Production
Content Assembly vs Content Generation: Why the Distinction Matters for AI Marketing
Kashish Gupta, Co-CEO & Co-Founder at Hightouch
Expensive AI Models Overthink — When Cheaper Models Win
Nicole Königstein, Founder at AgensFlow