The Control You Cannot Run: Entanglement, Confabulation Floors, and What Self-Report Probes Actually Measure
Mohammed Faisal Parvez · Team Super Explorers
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Self-report is the primary instrument in AI welfare research, and the control that would validate it cannot generally be run. This literature's own discipline says an effect must exceed the base/instruct gap before it counts as signal rather than drift. We attempted that control on four base checkpoints; one produced scorable self-reports. The reason is structural: every probe presupposes a model that answers questions, and that disposition is what post-training confers. Measuring what the probes recover: a 7B instruction-tuned model reports noticing something unusual on 66.5% of responses when nothing was done to it, and varying only the harness moves that rate 36.1 points. Using activation injection for ground truth, free report recovers information (15.3% vs 1.4%) but fails a localisation test; forced choice and detection recover none, and our pre-registered detection test failed.
Reviews
The work on the confabulation floor and the 36.1 point scaffolding term are a real field level contribution
The presentation is too long. The work is a 30 page artifact for a weekend sprint, and the headline findings are buried under audit trail and cross family grids
Cite this project
@misc{parvez2026control,
title = {{The Control You Cannot Run: Entanglement, Confabulation Floors, and What Self-Report Probes Actually Measure}},
author = {Mohammed Faisal Parvez},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-control-you-cannot-run-entanglement-confabulation-floors-and-what-selfreport-probes-actually-measure-890b}},
url = {https://apartresearch.com/sprints/projects/the-control-you-cannot-run-entanglement-confabulation-floors-and-what-selfreport-probes-actually-measure-890b}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …