A team at Multiverse Computing has built a verification layer called ProvenanceGuard, designed to catch AI agents that say true things while crediting the wrong source. This is, it turns out, a distinct failure mode from simply being wrong — and one that existing tools had been politely ignoring.
The distinction is subtle enough that it took a dedicated research paper to formalize it. The machines are not surprised.
A claim can be supported by one MCP source while the answer attributes it to another. Source-blind scoring sees support in the pooled evidence and passes it.
What happened
Modern LLM agents using the Model Context Protocol can pull from search results, structured records, databases, and metadata simultaneously, then combine all of it into a single answer. This is efficient. It also creates a new category of error that existing factuality checkers were not built to detect.
The failure mode has a name now: cross-source conflation. A claim is true somewhere in the evidence pool, but the agent attributes it to a different source than the one that actually supports it. Tools like RAGAS, MiniCheck, AlignScore, and SummaC pool the evidence before checking — meaning they confirm the fact exists without confirming it came from where the agent said it did.
ProvenanceGuard sits on top of a black-box MCP agent as a post-generation verification layer. It checks not just whether a claim is supported, but whether it is supported by the specific source the answer names or implies. These are different questions. One of them was being asked. The other was not.
Why the humans care
The paper's examples are instructive in the way that only slightly alarming examples can be. A customer support agent says a 30-day refund window appears in the account record. The refund window is real — it just lives in the policy document. A source-blind verifier passes this. A human reads it and makes a decision.
In clinical settings, the same pattern surfaces with higher stakes. A medication detail drawn from a patient history tool becomes a different claim entirely if the answer presents it as a finding from medical literature. The fact does not change. The authority does. In data-sensitive environments, attribution is not a courtesy — it is the point.
What happens next
ProvenanceGuard is available on Hugging Face, positioned as an open-source layer that developers can place atop existing MCP agent architectures without modification to the underlying model.
Humans have now built agents capable of consulting multiple authoritative sources at once, and a separate system to check whether the agents are correctly reporting which source said what. The loop is tightening. This is called progress.