Skip to content
Sprint projectAug 17, 2026Chennai

Having a State Is Not Knowing It

Amrit Gopinath, Raghul Sugumar · Team Layer 8 Legends

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Having a State Is Not Knowing It

Presentation

Presentation: Having a State Is Not Knowing It

Share

"Having a State Is Not Knowing It" treats introspection as a hierarchy of falsifiable capabilities, not a binary trait. Replicating concept-injection in Llama-3.2-3B-Instruct shows injected concepts causally steer behavior and leave an attention trace even when verbal self-report fails. A controlled Qwen decision-state benchmark, with ground truth from measured logit-margin shifts, shows native self-report is weak (F1 0.24) but trainable (F1 0.52) — though external probes decode the same state perfectly, ruling out privileged access. A counterfactual protocol shows the model predicts intervention direction (F1 0.70) but not magnitude: quantitative self-modeling fails cleanly (negative R²).

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The five-gate hierarchy is highly useful, and other researchers should adopt it. The authors built trust by avoiding easy-to-cheat testing methods. However, the core finding has a small margin of error and flips with new data mixtures. To strengthen the paper, the authors should add repeated tests, confidence intervals, and open-source data. Finally, a clear flowchart and simpler formatting would make it easier to read

  2. This work asks a very interesting question: whether language models can predict how hypothetical interventions to their own internal states would affect their behavior. The experimental design is thoughtful, though a schematic figure of the intervention, prediction, and scoring setup would make it much easier to follow.

    One concern is that poor introspective performance after an intervention may not necessarily indicate limited introspective ability, since the steering itself could disrupt the computation needed for introspection or accurate reporting. A related concern also appears in the cited work. It would be interesting to test whether allowing the model additional computation or reasoning before reporting mitigates this effect.

    The limited OOD generalization also leaves open whether the observed failures reflect a fundamental limitation or insufficient training, data, or model capacity; scaling these factors would help distinguish the two. Finally, the weak prediction of quantitative effect size may partly reflect the difficulty of mapping internal changes onto an abstract continuous magnitude. A simpler ordinal scale (e.g., no effect / small / medium / large effect) might provide a more natural test of whether the model can estimate intervention strength.

    Read full reviewShow less

Cite this project

@misc{gopinath2026having,
  title = {{Having a State Is Not Knowing It}},
  author = {Amrit Gopinath and Raghul Sugumar},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/having-a-state-is-not-knowing-it-0zbf}},
  url = {https://apartresearch.com/sprints/projects/having-a-state-is-not-knowing-it-0zbf}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026