Why Enterprise RAG Fails: The Document Layer Nobody Audits
Teams with bad retrieval reach for a new embedding model. The documents underneath rarely get checked. Six document-layer failure modes, and an audit for finding yours.

Retrieval is returning junk, so someone proposes a different embedding model. Two weeks later the chunks coming back are different junk. Then it is the vector store, then a reranker, then a quiet suggestion that maybe the base model is not good enough for this domain. Every one of those steps has a vendor attached, a benchmark to point at, and a story that sounds like engineering. Meanwhile nobody on the team has opened the documents the system is actually searching. In our experience that is where the problem has been sitting the entire time.
Key Takeaways
- Retrieval quality is capped by corpus quality. If the passage that answers the question was destroyed during ingestion, no embedding model will find it.
- The document layer has six parts: ingestion, chunking, metadata, versioning, permissions, and refresh. Each one can degrade retrieval on its own, and none of them throw an error when they do.
- In Anthropic's experiments, adding contextual information during indexing, alongside contextual BM25, reduced top-20 retrieval failures by 49% before reranking.
- Read the retrieved chunks, not the generated answers. Sampling 50 real questions and scoring whether a competent human could answer from the retrieved text alone is what separates a retrieval problem from a generation problem.
- A longer context window is not a substitute for precise retrieval. Model performance is often highest when the relevant information sits at either end of a long context, and can degrade when the model has to find it in the middle.
The reranker is not the problem
We have written before about whether your architecture can handle AI workloads and whether your software architecture is AI ready. Both of those are questions about compute, scale, and system design. This is a different question and it sits earlier in the chain: before you ask whether your architecture can serve retrieval at volume, it is worth asking whether the things being retrieved are worth serving.
Pavel Spesivtsev made this point on our podcast in an episode about knowledge infrastructure mattering more than models. After eighteen months of agentic implementation work, his read was that the models are rarely the weakest link: what breaks projects is incomplete specification, missing organizational knowledge, and poor context management. That is a claim about plumbing, not intelligence, and it is unglamorous enough that most teams skip past it on the way to the model selection debate.
This was the framing from the beginning. The original retrieval-augmented generation paper treated the corpus as the thing supplying knowledge and the generator as the part that phrases it. Somewhere between that paper and the current tooling market, the industry started treating the corpus as a solved input and the model as the interesting variable. For enterprise deployments where corpus quality has never been validated, that order of priorities deserves to be questioned.
What the document layer actually means
The document layer is everything between "a file exists somewhere in this organization" and "a passage of text arrives in a model's context window." It has six parts, and in a stalled project, any one of them can quietly undermine the entire retrieval pipeline:
- Ingestion and parsing. Getting bytes out of PDFs, Office files, wikis, ticketing systems, and scans, with structure intact.
- Chunking. Splitting documents into retrievable units without severing the context that made them meaningful.
- Metadata. Source, author, date, document type, business unit, jurisdiction, and anything else you will want to filter on later.
- Versioning. Knowing which of five similar documents is current and which was superseded.
- Permissions. Carrying access control from the source system into the index, so retrieval respects it.
- Refresh. Keeping the index in step with a corpus that changes after launch.
Only the first two are steps in the usual sense, run once as a document passes through. Metadata, versioning, permissions and refresh are standing conditions: they have to hold across the whole pipeline, and they are the ones that decay quietly after launch rather than failing on a particular day.
Most of this lives in the pipeline rather than the application, which is why it tends to fall between teams. Whoever built the ingestion and connector layer pulling from SharePoint, Drive, a document management system, or a claims platform is usually not the person tuning retrieval, and the failure surfaces in the second team's dashboard. Where that pipeline lives and who owns its quality is a software architecture decision made early and revisited rarely.
Six failure modes, and what each looks like from the outside
The useful thing about these is that they present as distinct symptoms. If you recognize the symptom, you can usually skip straight to the cause.
1. Structure-destroying parsing. The symptom is a system that is fluent and confidently wrong about anything that lived in a table. The cause is a text extractor that flattened a table into a run of space-separated values, with the header row now ten pages away from the numbers it labeled. Financial statements, rate schedules, benefits tables, and lab results all fail this way, and the failure is invisible unless someone reads the extracted text.
2. Chunks stripped of their context. The symptom is retrieved text that is topically relevant and practically useless: pronouns with no referent, "the company" with no company, "this policy" with no policy. The fix is upstream, and on at least one published evaluation it is measurable. Anthropic's work on contextual retrieval added a short generated description of each chunk's place in its parent document before embedding, and reported the top-20 retrieval failure rate falling from 5.7% to 2.9% when combined with contextual BM25, a 49% reduction. Adding a reranker on top took it to 1.9%. Those are their numbers on their corpora, not a figure to budget against. The transferable part is the ordering: the reranker helped, and it helped on top of the ingestion fix rather than instead of it.
3. Duplicates and superseded versions with no recency signal. The symptom is answers that sound correct and cite a policy retired two years ago. The cause is five near-identical documents in the index with nothing telling the retriever which one is live. This gets expensive fast in financial services and anywhere else the current version of a document is a compliance question rather than a preference.
4. Missing or wrong metadata. The symptom is that you cannot narrow anything: every query competes against the entire corpus, so a question about one business unit's procedure surfaces six other units' procedures with equal confidence. Metadata you did not capture at ingestion is expensive to reconstruct later.
5. Permissions that were never enforced at retrieval. This is the one that turns a quality problem into an incident. The symptom, if you are lucky, is a user seeing a snippet they should not have. If you are unlucky there is no symptom until someone outside the company notices. Carrying access control into the index as metadata is necessary but not sufficient. What matters is that retrieval authorizes against the requesting user's effective rights at the time of the request, and that it keeps doing so after permissions change in the source system, which they will. Pre-filtering the candidate set at query time and authorizing before content is returned are both defensible designs, and plenty of systems need both. What is not defensible is an index with no idea who is asking. This belongs in the security review rather than the backlog.
6. A corpus frozen at launch. The symptom is a system that was good at launch and is now vaguely disappointing, with no single change to blame. The cause is a one-time ingestion run and no refresh path. Quality did not break, it decayed, which is harder to notice and much harder to get funded.
The afternoon audit
Before buying anything, run this. A small team can start it in an afternoon. It gives you an initial picture of where retrieval is failing before you commit to another technology change.
- Collect 50 real questions. From production logs if you have them, from the people who will use the system if you do not. Questions you invented will be easier than real ones and will flatter the system.
- Run retrieval only. For each question, capture the chunks the system returns. Do not generate an answer. The answer is what has been hiding the problem.
- Score the chunks blind. Have someone who knows the domain read only the retrieved text, with the generated answer hidden, and mark each one: could a competent person answer this question correctly from this text alone? Yes, partly, or no. Separately, note whether anything else in the retrieved set contradicts the passage that answers it.
- Classify every failure into one of four buckets. The passage is not in the corpus at all. The passage is in the corpus but was not retrieved. The passage was retrieved but arrived mangled or stripped of context. Or the right passage was retrieved and arrived intact, but superseded or contradictory passages came with it, making it difficult for the generator to determine which source is authoritative.
Each bucket points at a different fix, and only the second is a retrieval tuning problem. The first is a coverage question about what you ingested. The third is parsing and chunking. The fourth is versioning and metadata, and it is the one a scorecard that only asks "did we retrieve it" will happily record as a pass, which is exactly why it survives in production so long.
Treat 50 questions as a diagnostic sample rather than a benchmark. It is enough to find the failure modes you have. It is not enough to publish a retrieval accuracy figure, and you should be suspicious of anyone who does so off a sample that size. This is the same exercise we run at the start of a software audit, deliberately cheap enough to do before anyone signs anything.
Fixing upstream beats tuning downstream
Roughly in order of leverage:
- Parse for structure, not just text. Tables should survive as tables, headings as headings, and document hierarchy as metadata. This is boring work with an unusually high ceiling.
- Chunk with context preserved. Carry section headings, document title, and effective date into the chunk, or generate a short contextual preamble the way the contextual retrieval work does.
- Make metadata first-class, and make authorization real. Any filter you will need at query time has to exist as an indexed field. Authorization has to resolve against the requesting user's current rights, whether you enforce that by pre-filtering the candidate set, by checking before content is returned, or both.
- Build a refresh path before launch. Decide how a changed document reaches the index, and how a deleted one leaves it. Retrofitting this is materially harder than building it.
To be clear about where the model does matter: a stronger model genuinely helps with synthesis across several retrieved passages, with following a required answer format, and with handling ambiguous questions. Those are real capability gains worth paying for. What it does not do is recover information that was destroyed before it arrived.
The same caution applies to context windows. The instinct once windows got large was to retrieve more and let the model sort it out. Research on long contexts found performance is often highest when the relevant information sits at the beginning or end of the input, and degrades significantly when the model has to find it in the middle, even for models built for long contexts. Retrieving twenty mediocre chunks instead of five good ones can make things worse.
When you actually do need the harder infrastructure
None of this is an argument that retrieval infrastructure is unnecessary. At genuine scale, across multimodal corpora, or where users actively probe for things they should not see, the heavier machinery earns its cost. Hybrid search, rerankers, query rewriting, and graph-augmented retrieval all solve real problems. The argument is about sequence. Buy them once you know your documents parse cleanly, your chunks carry their context, your metadata supports the filters you need, and your permissions hold. Buy them before that and you are paying to search a corpus you have not read.
Ankur Mattoo made a version of this point on the podcast when describing the AI foundations that made enterprise machine learning scalable: the organizations that moved quickly on generative AI were largely the ones that had done the unglamorous data work years earlier, for reasons that had nothing to do with generative AI. That foundation was not built for this moment. It just happened to be what made this moment possible.
If your retrieval quality has plateaued and every proposed fix involves buying something, the audit above is the cheaper first move. If you would rather someone else ran it, our AI practice does corpus and ingestion reviews as a scoped engagement, and the output is a list of what to fix in what order, whether or not we do the fixing.