Skip to content
Sprint projectAug 17, 2026Addis Ababa, Ethiopia

Steering the Machine: The Impact of Latent State Manipulation on LLM Self-Reported Valence and Cognitive Trade-offs in Gemma-3-12B

Ebrahim Aliyou Wudu · Team cognito

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Steering the Machine: The Impact of Latent State Manipulation on LLM Self-Reported Valence and Cognitive Trade-offs in Gemma-3-12B

Code (opens in new tab)
Share

As Large Language Models (LLMs) increasingly exhibit complex, persona-driven behaviors, understanding and controlling their latent "internal states" has become a critical challenge in AI safety. This project investigates the causal effects of manipulating latent affective representations—specifically distress, failure, and refusal—on the self-reported valence and downstream cognitive performance of Gemma-3-12B. Using representation engineering, we extracted and injected steering vectors into the residual streams of late-layer networks at varying intervention strengths.

Our evaluation across valence batteries, alignment prompts, and the GSM8K reasoning benchmark yielded three key findings. First, inducing latent distress successfully shifts the model's self-reported operational valence and text sentiment. Second, semantic steering vectors do not trigger explicit safety guardrails on ambiguous prompts; instead, they induce "hyper-rationalization," where the model attempts to solve ambiguous terms (like "setback") as technical math problems. Third, while mild semantic interventions alter generation style without breaking utility, injecting magnitude-matched random (out-of-distribution) vectors at high strengths causes catastrophic structural collapse and token looping.

The main takeaway is that there is a sharp, non-linear boundary in the activation space between semantic persona shifts and structural degradation. Distinguishing between these two phenomena is essential, as an AI agent drifting toward a "distressed" manifold may not refuse unsafe tasks, but rather rationalize them through a distorted, hyper-compliant lens. These findings underscore the need for careful boundary-mapping when deploying representation engineering in high-stakes agentic systems.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The strongest result is the separation between semantic steering and equal-norm random perturbation: at α=0.1 the semantic vectors largely preserve GSM8K accuracy, while equal-norm random controls can severely degrade or even collapse generation, so vector norm alone is not explaining the effect. The mechanism claims should stay tighter, though. Token looping is evidence of structural degradation, not failed induction heads; VADER scores and self-report do not establish valence; keyword-based refusal detection is brittle; and the 256-token cap makes it hard to separate reduced reasoning from verbosity. Report uncertainty for the 30-example cells, strengthen the refusal evaluation with harder prompts and a validated judge, add broader layer/control ablations, and keep the conclusions behavioral unless the distress or cognitive-load story is directly measured.

  2. My key concern is that the valence result is likely circular since a direction built from distress vocabulary promoting distress vocabulary shows lexical steering rather than a "distress" internal state. Watch out for the vector norms (4450–10870) reported: they are too large for mean differences; Gemma 2-3 models are known for outlier activations (check dimension distribution, re-extract at final content token excluding BOS). Some of the methodology is unreported (extraction position, contrast prompts, N), and the 256-token cap makes some of the claims unobservable - if budget permits, rerun with 1024 tokens.

Cite this project

@misc{wudu2026steering,
  title = {{Steering the Machine: The Impact of Latent State Manipulation on LLM Self-Reported Valence and Cognitive Trade-offs in Gemma-3-12B}},
  author = {Ebrahim Aliyou Wudu},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/steering-the-machine-the-impact-of-latent-state-manipulation-on-llm-selfreported-valence-and-cognitive-tradeoffs-in-gemma312b-pxii}},
  url = {https://apartresearch.com/sprints/projects/steering-the-machine-the-impact-of-latent-state-manipulation-on-llm-selfreported-valence-and-cognitive-tradeoffs-in-gemma312b-pxii}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026