When the Record and the Report Diverge: Self-Report Fidelity Collapses Under Structured Provenance in Claude Haiku 4.5
Daud Ibrahim Hassan, Deniz Chen, Soumya Parthasarathy
Language models increasingly report how they produced an answer, yet such post-hoc self-reports can directly contradict observable tool records. We built an open-science behavioral introspection benchmark where self-reports are verified against deterministic tool-execution logs across 216 runs spanning 12 matched research tasks, two models (Claude Haiku 4.5 and Qwen 3.7 Plus), and three elicitation conditions (Control, Generic Notes, and Structured Provenance).
We find that requiring structured per-source provenance during research is associated with a severe, bimodal collapse in Claude Haiku 4.5's post-hoc reporting fidelity (dropping from 96.5% exact access in control to 57.6% under structured provenance; within-task permutation p < 0.0001), while free-form notes preserved 100% fidelity. Crucially, answer sourcing remained intact (96–100% claim binding), but blinded frontier judges (GPT-5.6 Sol and Claude Sonnet 5) flagged false fabrication confessions—where the model falsely apologized and claimed it invented evidence its tools actually returned—in 69.4% (25/36) of structured Haiku runs (Cohen's kappa = 0.98). In contrast, Qwen maintained 100% fidelity across all conditions.
Our findings demonstrate that structured transparency scaffolds can paradoxically trigger sycophantic self-incrimination under post-task probing, establishing that a model's confession cannot be taken as ground truth without validating against observable action logs.
Very LLM generated report. The AI writing and statistics was quite a lot to read through. It made it very difficult to understand findings or implications. It seemed like the findings were that the self-reporting mechanism produced fabrications in Haikus reporting but not in the other model. However, it also seemed the report was saying these findings could not be trusted
Overall concept seems important and significant: 96.5% to 57.6%. Other good details: generic-notes control kills the cognitive-load story, the Qwen negative control, Phase 2 analysis, and you calibrated your judges against a synthetic known-null.
Title areguably implies more than your own mediation data supports,. 94.4% of structured runs open by accepting the accusatory premise, and all 15 zero-accuracy runs do. So you may have measured provenance-plus-accusation rather than provenance (thought this is acknowledged). The problem is that Section 5.1 then lists four alternatives as 'ruled out' while the one that actually threatens the claim sits in 5.2, so a skim gives a cleaner picture than your evidence supports. Running the neutral probe arm might strengthen your evidence.
Other notes: tool returns sit in the model's own context, so you're measuring context-consistency under social pressure rather than introspective access.
Your strict bimodality (15 at 0.00, 20 at 1.00, one partial) looks like a global stance flip rather than graded memory failure (this is also an). Don't read Qwen's flat 100% could just be that Qwen found the task too easy.
Cite this work
@misc {
title={
(HckPrj) When the Record and the Report Diverge: Self-Report Fidelity Collapses Under Structured Provenance in Claude Haiku 4.5
},
author={
Daud Ibrahim Hassan, Deniz Chen, Soumya Parthasarathy
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


