Validate the State Before Testing Introspection: A Causally Gated Protocol for Model Self-Report
Ngo Thai Bao
We developed a causally gated protocol for testing whether a language model has privileged access to an experimentally induced internal state. Using Qwen2.5-7B-Instruct, we extracted a candidate epistemic-deference activation direction, validated the intervention on development data, and required it to pass held-out behavioral and discriminant tests before any self-vs-observer introspection claim could be evaluated. The model showed strong intervention localization in its response distribution (Self AUC = 0.857, permutation p = 5×10⁻⁵) despite chance greedy A/B reports, and the candidate deference direction was highly stable across extraction splits (cosine = 0.956). However, on held-out behavioral families, the intervention increased concession probability by only 0.035 relative to a matched unrelated activation control and 0.071 relative to no intervention, below our preregistered minimum effect of 0.08. We therefore withheld the downstream privileged-introspection test. The main takeaway is methodological: statistically detectable or decodable activation changes should not be treated as valid introspection targets until their intended behavioral state has been independently and causally validated.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Validate the State Before Testing Introspection: A Causally Gated Protocol for Model Self-Report
},
author={
Ngo Thai Bao
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


