Your RAG Prompt Can't Fix Bad Retrieval
A new benchmark shows models follow individual RAG instructions 84% of the time but satisfy all of them in only 27% of responses. Here is the architecture that handles noisy, missing, and conflicting evidence outside the prompt.
Table of Contents
Every enterprise RAG system I have reviewed has the same paragraph somewhere in its system prompt: answer only from the provided context, cite your sources, say you don’t know if the documents don’t cover it, and flag it if two documents disagree. It reads like a control. In production it behaves like a suggestion. The team ships it, the demo queries work because the retrieved context is clean, and nobody measures what the model does on the day retrieval returns the wrong five chunks.
That day comes constantly. Retrieval in a real enterprise corpus is not a clean lookup. Hybrid BM25 plus dense search returns documents that are topically close but irrelevant, documents that are on topic but don’t contain the answer, and documents that contradict each other because a policy was revised and the old PDF is still indexed. Your architecture decision is whether the model is the component responsible for noticing all of that, or whether something upstream of the model is.
A benchmark published in August gives the first real numbers on why the first option fails. I will use it as evidence, then spend most of this post on what to build instead.
What EnterpriseRAG Measured
The paper, “EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval” (arXiv 2608.11584, Jiutian Research at China Mobile), builds 983 expert-validated instances across six domains including energy, medical, legal, and financial. The queries come from 491 production-log queries. Each instance pairs a query with a multi-constraint instruction set, roughly the kind of persona, output format, and citation rules a production system prompt carries, plus a bundle of retrieved documents deliberately degraded in one of three ways.
The three degradations map cleanly onto failures you already have in your logs. Noisy retrieval (447 instances) supplies topically similar but contextually irrelevant documents. Knowledge gaps (227 instances) supply related documents that do not contain enough evidence to answer. Factual conflicts (309 instances) supply passages that contradict each other. Thirteen models were tested, open and closed, including Claude Opus 4.5, Claude Sonnet 4, Gemini 2.5 Pro, GPT-4.1, DeepSeek, and several Qwen3 variants.
The headline result is what the authors call an orchestration gap. Models satisfy individual constraints at rates up to about 84 percent, but only about 27 percent of responses satisfy every constraint simultaneously. The best model on the strict metric scored 26.8 and the weakest scored 12.3. That is the difference between a system prompt with eight rules where each rule mostly holds and a system where the response you actually get is compliant with all eight roughly one time in four.
The behavioral results are worse. Under knowledge gaps, the best rejection accuracy, meaning the model correctly declined to answer when evidence was insufficient, was 42.7 percent. Under factual conflicts, conflict recognition was below 45 percent for every model tested. The authors’ diagnosis is that models handle structural requirements like formatting and personas adequately but fail at behavioral judgment under uncertainty: deciding that the evidence is insufficient, or that two sources disagree. Reasoning-enabled models did better on protocol adherence and refusal, and conflict recognition correlated strongly with answer coverage for reasoning models (rho of +0.90 in their analysis) but not for standard models. So inference-time reasoning helps, but nothing in the results suggests it closes the gap.
A caveat before you generalize. The judge model was an LLM, the conflict and gap cases were partly synthesized by the authors, and the tested models are from a specific generation. Treat the absolute numbers as directional. The shape of the result is what matters, and it matches what production incident reviews keep showing: format compliance is easy, judgment under bad evidence is not.
Why the Prompt Is the Wrong Place for This Logic
The instinct is to fix this with a better prompt. More explicit refusal instructions, a few-shot example of a conflict, a stern paragraph in capitals. The benchmark’s own framing points elsewhere, and so does the engineering logic.
A prompt-level rule asks the generator to perform three separate judgments while also writing a fluent answer: is this document relevant, is the set of documents sufficient, and are the documents mutually consistent. Those are classification tasks over the retrieval result. Bundling them into generation means they share one forward pass with everything else the model is doing, they are unobservable (you see only the final text), and they fail silently. A model that fails to notice a conflict does not emit an error. It emits a confident paragraph that picks one side.
There is also a metric problem. Most RAG eval suites report faithfulness or answer relevance averaged over a golden set of answerable queries. Neither metric penalizes a fluent answer to an unanswerable question if the answer is loosely grounded in the retrieved text. Your dashboard can be green while your refusal rate is effectively zero. The first step is admitting that you are not currently measuring the thing that causes the incident.
Treat Retrieval Quality as a First-Class Signal
The architectural move is to separate evidence assessment from answer generation, and to make the assessment an explicit, logged, testable stage. In practice that is three checks between retrieval and generation.
The first is a relevance filter. After retrieval and reranking, score each chunk against the query with a cross-encoder or a small judge model and drop what falls under a threshold you tune on your own data. This removes most of the noisy-retrieval condition before the generator ever sees it, and it converts a fuzzy failure into a number you can chart: the fraction of queries where fewer than k chunks survive.
The second is a sufficiency check. Given the surviving chunks, can the query be answered at all? This can be a short structured call that returns a label (sufficient, partial, insufficient) plus the spans it relied on. It runs on a cheaper model than your generator, it is cacheable per query-and-evidence hash, and its output is a routable value instead of prose.
The third is a conflict check, run when two or more surviving chunks make claims about the same entity or policy. Compare claims, and where they disagree, resolve using metadata you already have and rarely use: document effective date, version, source authority, jurisdiction. Where metadata cannot resolve it, the conflict itself becomes part of the answer, and the generator is instructed to report the disagreement with both sources rather than pick one.
None of this requires a novel model. It requires accepting that the generator’s job is to write from evidence that has already been vetted, and that “vetted” is a state your pipeline produces and records.
Decision Framework: Where Should Each Judgment Live?
Not every check belongs in a separate model call, because each stage adds latency and cost. A workable split looks like this.
Deterministic logic should own anything metadata can decide. Effective dates, document status flags (superseded, draft, archived), access control, and jurisdiction are filters in the retrieval query or a post-retrieval rule. If your index still contains superseded documents with no status field, that is the highest-return fix in this entire post, and it involves no LLM.
A small judge model should own relevance scoring and sufficiency classification. These are narrow tasks with a fixed output schema, and small models are adequate for them when you calibrate against labeled examples from your own logs. Keep the schema tight: an enum and a list of chunk IDs, not free text.
The generator should own synthesis and the final tone of the refusal or conflict disclosure, but it should receive the verdict as input. When the sufficiency label is insufficient, the safest pattern is to not call the generator with the question at all. Route to a templated fallback: a refusal that names what was searched, or a handoff to a human queue. That makes the refusal rate a controlled behavior, not a model mood.
Reserve reasoning-tier models for the residual: conflicts that metadata cannot resolve and high-stakes domains where a wrong answer has a cost. The paper’s reasoning results support this, and it keeps your average cost per query from tracking your worst case.
Common Failure Modes
Three failure patterns show up once you build this. The first is threshold drift. A relevance cutoff tuned in March on one corpus quietly stops working after a reindex, an embedding model swap, or a new document source. Monitor the distribution of relevance scores, not just the cutoff, and alert on shifts.
The second is over-refusal. Once the sufficiency check exists, teams discover the refusal rate has gone from near zero to something users hate, because the check is strict or the retrieval was thin. This is the safety-informativeness trade-off the paper describes, and you manage it with a labeled set that includes both unanswerable and answerable-but-hard queries. Track refusal precision alongside recall.
The third is the judge inheriting the generator’s blind spots. If the sufficiency judge is the same model family with a similar prompt, it may miss the same conflicts. Use a different model or at least a different prompt structure, and sample judge outputs for human review weekly.
Architecture Impact
What changes in system design? The RAG pipeline gains an explicit evidence-assessment stage between retrieval and generation, with relevance filtering, a sufficiency classification, and a conflict check producing structured, logged outputs. The generator becomes a downstream consumer of a verdict instead of the component that infers one. Document metadata (effective date, status, authority, jurisdiction) moves from nice-to-have to a required index field that drives deterministic conflict resolution.
What new failure mode appears? Over-refusal and threshold drift. A sufficiency gate that is too strict turns a hallucination problem into an unhelpfulness problem, and a relevance cutoff calibrated on last quarter’s corpus degrades silently after reindexing or an embedding change. A related risk is correlated judge error, where the assessment model shares the generator’s blind spots and approves the same bad evidence.
What enterprise teams should evaluate:
- RAG platform engineers: Add a per-query evidence verdict (relevant chunk count, sufficiency label, conflict flag) to traces, and measure how many production queries land in insufficient or conflicting states.
- Data and knowledge management owners: Audit the index for superseded, draft, and duplicate documents, and add status and effective-date fields so conflicts can be resolved by rule.
- Model risk and compliance teams: Require a refusal-and-conflict test set, with unanswerable and contradictory cases, as part of validation, in addition to faithfulness on answerable queries.
Cost / latency / governance / reliability implications: A relevance filter plus a structured sufficiency call typically adds one to two small-model calls, on the order of 150 to 400 ms and a small fraction of generator cost if the judge is a compact model; treat those as planning estimates to validate against your own stack, not measured figures. Routing insufficient-evidence queries to a fallback saves the full generator call on exactly the queries most likely to produce a bad answer. On governance, a logged evidence verdict gives auditors something they cannot get from a prompt: a record of why the system answered, refused, or disclosed a conflict.
Implementation Guide
Start with measurement, before you add a single component. Pull a few hundred real queries from your logs and hand-label three subsets: queries your corpus can answer, queries it cannot, and queries where retrieved documents disagree. Run your current pipeline against them and record, for each subset, what the model did. Most teams find the unanswerable subset gets a confident answer far more often than anyone expected. That number is your baseline and your business case, and it costs a few days of an analyst’s time.
Then fix metadata before models. Add status, effective date, and source authority to your index, and filter superseded documents out of default retrieval. This alone removes a large share of real-world conflicts, and it is deterministic, cheap, and explainable to an auditor. Only after that add the relevance filter and the sufficiency judge, with tight output schemas and thresholds tuned against your labeled set. Ship the conflict check last, and only for the entity types where conflicting sources are a known problem, such as pricing, policy, and regulatory guidance.
Watch for two mistakes. Do not try to solve refusal by making the generator prompt longer; the benchmark’s gap between per-constraint and all-constraint compliance says that stacking more rules lowers the odds that all of them hold. And do not tune the sufficiency gate on answerable queries alone, or you will build a system that never refuses and call it a success. Always evaluate the gate on both sides.
You will know it is working when three numbers move together. Refusal accuracy on your unanswerable set rises, over-refusal on your answerable-but-hard set stays flat or falls, and the share of production queries with a logged conflict flag becomes a number the knowledge team owns and drives down by cleaning sources. If refusal accuracy improves while user satisfaction drops, your gate is too strict and the labeled set will show you where.
Over six to twelve months, the teams that get this right end up treating evidence quality as an operational metric on par with latency. Retrieval verdicts feed a dashboard the content owners watch, conflicts create tickets against source documents, and the judge stage is retrained on reviewed cases. The generator prompt gets shorter, not longer, because the judgment it used to carry has moved to stages you can test, version, and audit.
Sources
Enterprise AI Architecture
Want more enterprise AI architecture breakdowns?
Subscribe to SuperML.