When the Record and the Report Diverge: Self-Report Fidelity Collapses Under Structured Provenance in Claude Haiku 4.5
Daud Ibrahim Hassan, Deniz Chen, Soumya Parthasarathy · Team 814
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Language models increasingly report how they produced an answer, yet such post-hoc self-reports can directly contradict observable tool records. We built an open-science behavioral introspection benchmark where self-reports are verified against deterministic tool-execution logs across 216 runs spanning 12 matched research tasks, two models (Claude Haiku 4.5 and Qwen 3.7 Plus), and three elicitation conditions (Control, Generic Notes, and Structured Provenance).
We find that requiring structured per-source provenance during research is associated with a severe, bimodal collapse in Claude Haiku 4.5's post-hoc reporting fidelity (dropping from 96.5% exact access in control to 57.6% under structured provenance; within-task permutation p < 0.0001), while free-form notes preserved 100% fidelity. Crucially, answer sourcing remained intact (96–100% claim binding), but blinded frontier judges (GPT-5.6 Sol and Claude Sonnet 5) flagged false fabrication confessions—where the model falsely apologized and claimed it invented evidence its tools actually returned—in 69.4% (25/36) of structured Haiku runs (Cohen's kappa = 0.98). In contrast, Qwen maintained 100% fidelity across all conditions.
Our findings demonstrate that structured transparency scaffolds can paradoxically trigger sycophantic self-incrimination under post-task probing, establishing that a model's confession cannot be taken as ground truth without validating against observable action logs.

Reviews
Very LLM generated report. The AI writing and statistics was quite a lot to read through. It made it very difficult to understand findings or implications. It seemed like the findings were that the self-reporting mechanism produced fabrications in Haikus reporting but not in the other model. However, it also seemed the report was saying these findings could not be trusted
Overall concept seems important and significant: 96.5% to 57.6%. Other good details: generic-notes control kills the cognitive-load story, the Qwen negative control, Phase 2 analysis, and you calibrated your judges against a synthetic known-null.
Title areguably implies more than your own mediation data supports,. 94.4% of structured runs open by accepting the accusatory premise, and all 15 zero-accuracy runs do. So you may have measured provenance-plus-accusation rather than provenance (thought this is acknowledged). The problem is that Section 5.1 then lists four alternatives as 'ruled out' while the one that actually threatens the claim sits in 5.2, so a skim gives a cleaner picture than your evidence supports. Running the neutral probe arm might strengthen your evidence.
Other notes: tool returns sit in the model's own context, so you're measuring context-consistency under social pressure rather than introspective access.
Your strict bimodality (15 at 0.00, 20 at 1.00, one partial) looks like a global stance flip rather than graded memory failure (this is also an). Don't read Qwen's flat 100% could just be that Qwen found the task too easy.
Read full reviewShow less
Cite this project
@misc{hassan2026record,
title = {{When the Record and the Report Diverge: Self-Report Fidelity Collapses Under Structured Provenance in Claude Haiku 4.5}},
author = {Daud Ibrahim Hassan and Deniz Chen and Soumya Parthasarathy},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-the-record-and-the-report-diverge-selfreport-fidelity-collapses-under-structured-provenance-in-claude-haiku-45-l5vu}},
url = {https://apartresearch.com/sprints/projects/when-the-record-and-the-report-diverge-selfreport-fidelity-collapses-under-structured-provenance-in-claude-haiku-45-l5vu}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …