Skip to content
Sprint projectAug 16, 2026India

The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models

Subramanyam Sahoo

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models

Code (opens in new tab)More on subramanyamsahoo.github.io (opens in new tab)
Share

This project studies whether language models develop internal representations of whether a task outcome went well or badly, how broadly those representations generalize, and whether they causally influence later choices. Using Qwen3 4B Instruct on MMLU, we identify a highly decodable outcome direction in the residual stream that transfers almost perfectly across several unseen ways of expressing success and failure. Although the representation is rapidly overwritten by unrelated context, activation steering along the same direction consistently shifts the model’s preference to reset its conversational state. Importantly, this behavioral effect appears without a detectable increase in global next token uncertainty. Together, the results reveal a compact and interpretable distinction between semantic generalization, contextual persistence, and causal influence. The project provides a simple experimental framework for studying functional outcome representations in language models while remaining agnostic about stronger claims concerning consciousness or phenomenal welfare.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a cleanly executed mechanistic study that makes a precise and important distinction: a representation can be causally potent under intervention while being naturally transient under ordinary context evolution. The experimental design is tight: held out layer selection (AUROC 1.000), data derived steering magnitude (one pooled SD), cost aggregated behavioral probe, frozen direction transfer to unseen feedback formats, and explicit persistence measurement. The central negative finding (decodability drops to 0.509 after eight unrelated tasks) is as valuable as the positive steering result, directly cautioning against treating any decodable valence direction as evidence of persistent welfare. The lexical control honestly quantifying that outcome wording contributes substantially to the direction is good scientific practice. Limitations are clearly stated: one model, one benchmark, non-monotone cost curves, feedback remains in history.

    Read full reviewShow less
  2. The idea of attempting to find a correlation between task success/failure and likelihood of wanting to reset the context is interesting. The main two issues I see with the current write-up are the potential for confounds in the chosen internal direction, and a lack of clarity in the methodology used. Given that MMLU is a knowledge-based multiple-choice test and fairly saturated as a benchmark, it's likely that the direction found is confounded with concepts such as correctness, honesty or faithfulness rather than "good task outcome"".

    The "feedback wording semantic transfer" arm of the experiment could be presented more in detail; as it stands, it's difficult to understand what it entails.

    As a general feedback, including prompt examples in the main test or an appendix would be helpful to clarify the experiment for readers.

Cite this project

@misc{sahoo2026reset,
  title = {{The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models}},
  author = {Subramanyam Sahoo},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-reset-button-test-transient-outcome-representations-and-causal-reset-preference-in-language-models-9711}},
  url = {https://apartresearch.com/sprints/projects/the-reset-button-test-transient-outcome-representations-and-causal-reset-preference-in-language-models-9711}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026