Skip to content
11 min read

The Right Fact From the Wrong Source Is Still a Wrong Answer

Enterprise assistants now answer from several systems at once. A claim can be true somewhere in that evidence and still credit the wrong record, and pooled checks will pass it.

Antonio J. del Águila

Knaisoma

Consider a support assistant that answers a customer with this sentence: “According to your account record, your plan includes a 30-day refund window.” Every word of the refund policy is accurate. The general policy document says exactly that. The account record, however, says nothing about refunds, and this customer signed an enterprise agreement whose terms differ. The assistant has produced a true fact, a real citation, and a wrong answer, because the reader will act on the claim that the policy applies to them specifically.

The example is adapted from a September 29 write-up by researchers at Multiverse Computing, who call this failure cross-source conflation: a claim that is true somewhere in the evidence but attributed to the wrong source. What has changed is the architecture organizations are deploying. Assistants for support, legal review, finance operations and internal knowledge no longer answer from one document. They call several tools, often through the Model Context Protocol, and merge a customer record, a policy library, a ticket history and a contract repository into one context. Every additional source is another way to credit the right fact to the wrong place.

Our position is that for any assistant drawing on more than one system of record, attribution deserves its own check, separate from factual support, and the pipeline must be built so that check is possible. Most evaluation stacks today cannot perform it, because they discard the information it needs before the check ever runs.

Why a true sentence can be the failure

The clearest evidence that misattribution is a distinct and dangerous failure comes from law, where citations carry consequences. In 2024, a Stanford team including Varun Magesh and Daniel Ho evaluated the leading AI legal research tools from LexisNexis and Thomson Reuters, products built on retrieval and marketed partly on the reliability of their citations. The researchers separated two questions that most product evaluations merge: whether a response is correct, and whether it is grounded, meaning that its key propositions cite sources that actually support them. They defined a response as misgrounded when it cites a source that does not support the claim or that is inapplicable, and they counted misgrounded answers as hallucinations even when the statement itself was right.

Under that definition, the tools hallucinated on 17 to 33 percent of the benchmark queries. The tested versions are now more than two years old, so the specific rates describe those products at that time, not today. The reasoning has aged better than the numbers. The authors argued that these errors are potentially more dangerous than an invented case, because they are subtler and harder to spot. A fabricated citation fails the moment someone looks it up. A real citation attached to the wrong proposition survives the lookup, and only someone who reads the source and compares it with the claim will catch it.

Outside law, the Tow Center for Digital Journalism tested eight AI search tools in March 2025 by giving each one excerpts from real news articles and asking it to identify the source. Collectively, the tools answered more than 60 percent of the 1,600 queries incorrectly. DeepSeek credited the excerpt to the wrong source in 115 of 200 attempts, and several tools cited syndicated or copied versions of an article rather than the original publisher. It measured consumer search products, which have changed since, but it isolates the attribution question: systems that retrieve the right text often lose track of where it came from.

Internally, the cost of misattribution is less visible and often larger. An answer that credits a refund window to the account record changes what an agent promises a customer. A contract summary that cites the master agreement for a clause that lives in an expired amendment changes a negotiating position. An HR assistant that applies one country’s leave policy to an employee in another produces an answer that is correct for somebody. In each case the fact passes review because it is true, and the attribution is the part that determines whether it applies.

Where pooled verification goes blind

Teams that evaluate retrieval-augmented systems commonly measure faithfulness: whether each claim in an answer is supported by the retrieved context. Tools such as RAGAS, MiniCheck and AlignScore are useful, and a team without any of them should start there. In their usual configuration, though, they treat the retrieved context as one pool of evidence. The Multiverse authors make this point directly: the standard checkers ask whether a claim is supported once the evidence has been pooled, not which tool output supports it or whether that is the output the answer names. A misattributed claim is, by construction, supported somewhere in the pool, so a pooled checker passes it.

Their system, ProvenanceGuard, keeps source identity through verification. It decomposes an answer into claims, routes each claim to the source most likely to support it, checks support with a natural language inference model, and compares the source that actually supports the claim with the source the answer credits. On 361 claims from 40 held-out answers produced by a medical agent and reviewed by domain experts, it blocked 138 of the 139 claims the experts said should not pass. When the researchers changed the named source in 50 test cases while leaving the evidence intact, it caught all 50 swaps. Its block-decision F1 score of 0.802 was close to MiniCheck’s 0.783 and RAGAS faithfulness at 0.758, which tells you something important: on overall support, the pooled checkers are competitive. What they cannot do is say which source supports each claim, and in the authors’ comparison table ProvenanceGuard is the only verifier that emits that mapping.

Treat these numbers as one team’s research result rather than a product benchmark. The evaluation covers a single medical deployment, the authors are a vendor with an interest in the approach, and we found no independent replication of the results. Its limitations are also stated plainly, which makes it more useful than most launch material. When several sources looked alike, the system still blocked well, with an F1 of 0.846, but it identified the exact supporting source in only 50.3 percent of claims. It also held 67 claims the experts considered supported, sending them to review or repair.

Four questions every claim has to answer

The useful idea in this research is not a particular verifier. It is a change in the unit of evaluation, from the answer to the claim, and from “is this supported” to a short sequence of questions that each claim must pass. We find it helpful to write them down as the fields of a verification record, because a record makes it obvious which questions a pipeline can actually answer. The example below is an illustrative scenario built on the refund case above, not data from any deployment.

# One record per atomic claim, checked before release.
claim: "Your plan includes a 30-day refund window."

# 1. What does the answer say it relied on?
credited_source: crm.account/ACME-0042

# 2. What actually supports the claim?
supporting_source: policy.refunds/v7
support: entailed        # pooled faithfulness stops here

# 3. Is the credited source the supporting one?
attribution_match: false

# 4. Does the supporting source govern THIS reader?
authority:
  scope: general_policy
  overridden_by: contract.msa/ACME-0042
  version_current: true

decision: hold_for_review   # block, hold, repair, release
reason: "Policy text credited to the account record;
  the customer contract governs."

The first two questions require only that source identifiers survive the pipeline, which is harder than it sounds. The third question is the one pooled verification cannot answer, and it is where ProvenanceGuard adds its contribution. The fourth question goes further than the research does, and it is where most of the business risk lives. A claim can be correctly attributed to a source that has no authority over the reader’s situation: a superseded policy version, the wrong jurisdiction, a different customer’s record, a draft rather than an approved document. The Stanford team’s definition of misgrounding already includes the inapplicable source, and enterprise systems need the same rule, expressed in terms of their own data: which record governs this customer, which version is in force, which policy outranks which.

The fourth question cannot be answered by a language model alone. It needs metadata the organization already has but rarely passes to the assistant: document status, effective dates, precedence between contract and policy, the tenant or customer that a record belongs to. This is why attribution checking is partly a data engineering task, and why buying a better verifier will not close the gap on its own.

Keep source identity from leaking out upstream

Many pipelines lose attribution long before any check runs. A retrieval layer returns chunks without stable document identifiers. An orchestration framework concatenates tool outputs into a single block of text. A prompt template asks the model to answer and “cite sources” but supplies no identifiers to cite, so the model cites whatever looks plausible. By the time an evaluator sees the answer, the information needed to answer the second question has been thrown away.

The protocols themselves do not force this loss. The current Model Context Protocol specification for tools lets a tool return resource links and embedded resources that carry a URI, and every content type can carry annotations, including a last-modified time. That gives a server a standard place to state where each piece of content came from and when it changed. Nothing in the protocol requires a client to keep that information when it assembles a model’s context, and many integrations flatten everything into text. The engineering work is therefore unglamorous and concrete: keep a source identifier on every retrieved unit, keep it through every transformation, require the model to cite identifiers rather than titles, and log the mapping so a reviewer or verifier can reconstruct it.

Several recognizable antipatterns undo this work. The most common is a list of links at the bottom of an answer, which looks like attribution but cannot say which sentence came from which source. Another is an evaluation set built mostly from questions that a single document can answer, which never exercises conflation at all. A third is treating the presence of a citation as evidence of verification, which is precisely the reasoning the Stanford study took apart. The last is measuring a verifier by how much it blocks, when the cost that matters is how many correct answers it sends to a human queue.

What this costs, and when it is worth paying

Claim-level attribution checking is not free, and the research is candid about where the costs fall. The reported latency was roughly half a second per answer on a local configuration, modest for an offline gate or an asynchronous review step, and more noticeable in a live chat. The larger cost is review load. A verifier that holds 67 supported claims alongside 138 unsupported ones is conservative by design, and someone has to clear that queue. If the reviewers are the same subject matter experts the assistant was meant to relieve, an aggressive gate can consume most of the time the assistant saved. And because identifying the exact source fell to roughly half when sources looked alike, the hardest cases (several versions of the same policy, near-identical contracts, records of related customers) are where an automated check is weakest and human judgment is still needed.

That suggests a simple way to decide where to invest. Attribution checking earns its cost where three conditions hold together: the answer draws on more than one source, the sources disagree or apply to different parties, and the reader will act on which source said it. Record-specific answers about a customer, patient, employee or contract meet all three, and so do regulated answers where authority depends on jurisdiction or version. An internal assistant answering general questions from a single, current handbook meets none of them, and pooled faithfulness plus sampling review is a reasonable stopping point. Summarizing one document the user supplied cannot conflate sources at all. Between those ends sit internal knowledge assistants over versioned policies, where the cheapest effective step is usually metadata: carry status and effective dates into the context, and let a rule reject answers that cite superseded material before any model-based verifier runs.

Further reading: our earlier pieces on why observability is not evaluation and on the control plane that Model Context Protocol deployments need cover the measurement and governance layers this check depends on. The Stanford team’s full paper is worth reading for its typology alone.

If your assistants answer from customer records, contracts or policies held in several systems, attribution is probably the risk your current evaluations are not measuring. We help teams trace how source identity flows through retrieval and tool pipelines, design attribution test sets from their own documents, and decide where a verification gate is worth its review cost. Talk with us about grounding your AI assistants.

AI Agentic AI Evaluation Governance
Share:

Stay updated

Get insights on engineering transformation delivered to your inbox.

Newsletter coming soon.