Skip to content
Sprint projectAug 17, 2026Seoul, Korea

Which Way Was I Steered? Testing Signed Introspection in Gemma 3

Aiden Geunhyoung Kim · Team Meow

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Which Way Was I Steered? Testing Signed Introspection in Gemma 3

Share

Can a language model tell not only that its activations were perturbed, but which semantic direction they were moved? We test this in Gemma 3 27B using a bipolar positive–negative sentiment axis, equal-magnitude +v/−v interventions, counterbalanced forced-choice token-logit readouts, and a prior-only protocol in which steering is removed before the model answers. The axis was strongly separable (held-out AUROC 0.979), causally steered sentiment in both directions under continuous injection, and preserved control-task accuracy. However, after the intervention was removed, the model could neither detect nor identify its direction: pooled sign accuracy was 49.4% over 312 trials, while +v and −v detection were 51.0% and 52.1%, all at chance. This demonstrates a boundary between causal steerability and reportable introspective access: changing an internal representation does not imply that the model can later report the change.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The distinction you're drawing is a good one. Noticing that something was injected isn't the same as knowing which way; a detector that only measures distance from normal fires either way and tells you nothing about direction. Two opposite pushes of equal size is a clean way to force the harder question, since that kind of detector can't pass by construction. The engineering is careful for 24 hours: you reproduce the existing detection result before running anything new, use forced choice instead of yes/no to dodge the known bias toward "yes," and add a runtime check that kills the run if the steering hook is still attached at question time. The safety framing lands too; a self-report channel that only works while something is actively happening is useless at exactly the moment you'd want to query it.

    My main concern is that there's no positive control for the sign task itself. You ask a new question in a new format and get a chance. That's consistent with your conclusion, but equally consistent with the model being unable to answer this question in this format even when the answer is plainly available. Your ceiling comparison is the concept-detection replication, which you rightly flag as unmatched. The missing condition is the sign question with steering still on, same axis; if the model can report direction live but not after, the claim is locked down - you list this in future work. Related: your validation gates all run with steering active while the test runs with it removed, and your own "what didn't work" section says the removed condition showed no directional effect. Two small things. Your dose explanation for the disagreement with Pearson-Vogel is reasonable, but you differ on model, axis type, layer count, and protocol at once, so dose is one candidate of four. Housekeeping: your PDF title doesn't match the review form, and your LLM usage statement appears to have the abstract pasted into it.

    Read full reviewShow less
  2. The three validation checks make this null result convincing, and the equal-magnitude design is a good test.

    However, the missing part is the hidden-state analysis.

    Since the hidden states are already available, run this analysis to determine whether the internal trace is truly gone or just difficult to detect.

    Do this before publishing.

Cite this project

@misc{kim2026which,
  title = {{Which Way Was I Steered? Testing Signed Introspection in Gemma 3}},
  author = {Aiden Geunhyoung Kim},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/which-way-was-i-steered-testing-signed-introspection-in-gemma-3-l7cs}},
  url = {https://apartresearch.com/sprints/projects/which-way-was-i-steered-testing-signed-introspection-in-gemma-3-l7cs}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026