The Control You Cannot Run: Entanglement, Confabulation Floors, and What Self-Report Probes Actually Measure
Mohammed Faisal Parvez
Self-report is the primary instrument in AI welfare research, and the control that would validate it cannot generally be run. This literature's own discipline says an effect must exceed the base/instruct gap before it counts as signal rather than drift. We attempted that control on four base checkpoints; one produced scorable self-reports. The reason is structural: every probe presupposes a model that answers questions, and that disposition is what post-training confers. Measuring what the probes recover: a 7B instruction-tuned model reports noticing something unusual on 66.5% of responses when nothing was done to it, and varying only the harness moves that rate 36.1 points. Using activation injection for ground truth, free report recovers information (15.3% vs 1.4%) but fails a localisation test; forced choice and detection recover none, and our pre-registered detection test failed.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Control You Cannot Run: Entanglement, Confabulation Floors, and What Self-Report Probes Actually Measure
},
author={
Mohammed Faisal Parvez
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


