When AI Gets the Fact Right but Credits the Wrong Source, a New Tool Catches It
ProvenanceGuard targets a blind spot in AI fact-checking: a claim can be true and still be falsely attributed. In medical tests, it caught 138 of 139 claims that should have been blocked.

Key points
- ProvenanceGuard caught 138 of 139 claims that human experts said should be blocked, across 361 hand-verified examples drawn from medical AI agent outputs.
- The tool addresses "cross-source conflation," a failure where an AI's answer is factually correct but credits the wrong source, which standard fact-checkers miss entirely.
- On the primary held-out test, ProvenanceGuard scored 0.802 on the metric that balances catching bad claims against blocking good ones, beating four rival checkers including MiniCheck (0.783) and RAGAS Faithfulness (0.758).
- It runs in roughly half a second per answer in the tested setup, making it practical as an offline review gate.
Imagine asking an AI assistant about your health plan and getting back a confident answer: "According to your patient record, this medication is recommended for your condition." The detail might actually appear in a research article the AI also read, not in your personal record at all. The fact is real, the credit is wrong, and almost every tool designed to catch AI errors would pass it anyway.
That is the problem a new system called ProvenanceGuard is built to fix. Researchers published the work this week on arXiv, and it was also highlighted by Hugging Face.
Why does mixing up sources matter?
It matters most when the source carries legal or clinical weight. A wrong attribution is not just sloppy; it can send a clinician or a patient to the wrong place for verification, with consequences a simple factual error might not carry.
AI assistants increasingly work by pulling information from several places at once, using a standard called the Model Context Protocol, or MCP, which lets an AI call a search tool, a database, a patient record system and other sources in one go, then stitch the results into a single answer. Most fact-checking tools pool all that evidence together before testing whether the answer holds up. Pool the sources and the wrong-attribution problem disappears from view: the fact is supported somewhere, so the checker passes it.
The researchers call this failure mode "cross-source conflation." We first covered MCP-related agent behaviour on 14 July 2026, and our story on Google's smart-home agent integration last week showed how quickly these multi-source pipelines are reaching ordinary users.
| Verifier | Block/reject F1 score | Tracks which source? |
|---|---|---|
| ProvenanceGuard | 0.802 | Yes |
| MiniCheck | 0.783 | No |
| RAGAS Faithfulness | 0.758 | No |
| AlignScore | 0.662 | No |
| SummaC-ZS | 0.436 | No |
The F1 score here is a single number, ranging from 0 to 1, that captures how well a system catches bad claims without over-blocking good ones. Higher is better.
How does ProvenanceGuard actually work?
It sits on top of an existing AI agent and checks answers after they are produced, without retraining anything. Each source stays labelled separately throughout. The system then runs five steps: breaking the answer into individual claims, finding the most relevant source for each, checking whether that source actually supports the claim, comparing it to whichever source the answer names, and issuing a verdict to allow or block the answer.
Blocked answers do not simply disappear. A repair step, based on a technique called RARR (Retrofit Attribution using Research and Revision, originally described here), attempts a rewrite grounded in the correct source. When no trustworthy rewrite is possible, the system outputs a safe fallback rather than guess.
In the full test run of 173 blocked answers, all were resolved, though 144 ended as cautious fallback text. The researchers treat that as the system working correctly: a non-answer beats a confidently wrong one.
What should ordinary users take away?
If you use an AI tool that pulls from several data sources, such as a health app, a customer service bot, or a financial assistant, the system as currently built probably cannot tell you which source backs each thing it says. ProvenanceGuard is a step toward tools that can.
It's a research paper, not a finished product. But catching 138 of 139 claims that experts said should fail, while picking the right source 86% of the time, puts it meaningfully ahead of current standard tools on the specific problem it was designed to solve.
The harder challenge, telling apart two very similar sources, remains open: source identification accuracy dropped to 50.3% when sources closely resembled each other. That caveat matters in clinical settings, where two studies on the same drug can look nearly identical to an automated system. Our NHS watchdog story from 31 August showed what happens when clinical AI errors go unchecked at the source level.
The attribution problem ProvenanceGuard targets is one the field has largely been ignoring while arguing about other kinds of AI error. Getting the source right, not just the fact, is where the next reliability fight is.



