Having a State Is Not Knowing It
Amrit Gopinath, Raghul Sugumar
"Having a State Is Not Knowing It" treats introspection as a hierarchy of falsifiable capabilities, not a binary trait. Replicating concept-injection in Llama-3.2-3B-Instruct shows injected concepts causally steer behavior and leave an attention trace even when verbal self-report fails. A controlled Qwen decision-state benchmark, with ground truth from measured logit-margin shifts, shows native self-report is weak (F1 0.24) but trainable (F1 0.52) — though external probes decode the same state perfectly, ruling out privileged access. A counterfactual protocol shows the model predicts intervention direction (F1 0.70) but not magnitude: quantitative self-modeling fails cleanly (negative R²).
The five-gate hierarchy is highly useful, and other researchers should adopt it. The authors built trust by avoiding easy-to-cheat testing methods. However, the core finding has a small margin of error and flips with new data mixtures. To strengthen the paper, the authors should add repeated tests, confidence intervals, and open-source data. Finally, a clear flowchart and simpler formatting would make it easier to read
Cite this work
@misc {
title={
(HckPrj) Having a State Is Not Knowing It
},
author={
Amrit Gopinath, Raghul Sugumar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


