The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models
Subramanyam Sahoo
This project studies whether language models develop internal representations of whether a task outcome went well or badly, how broadly those representations generalize, and whether they causally influence later choices. Using Qwen3 4B Instruct on MMLU, we identify a highly decodable outcome direction in the residual stream that transfers almost perfectly across several unseen ways of expressing success and failure. Although the representation is rapidly overwritten by unrelated context, activation steering along the same direction consistently shifts the model’s preference to reset its conversational state. Importantly, this behavioral effect appears without a detectable increase in global next token uncertainty. Together, the results reveal a compact and interpretable distinction between semantic generalization, contextual persistence, and causal influence. The project provides a simple experimental framework for studying functional outcome representations in language models while remaining agnostic about stronger claims concerning consciousness or phenomenal welfare.
This is a cleanly executed mechanistic study that makes a precise and important distinction: a representation can be causally potent under intervention while being naturally transient under ordinary context evolution. The experimental design is tight: held out layer selection (AUROC 1.000), data derived steering magnitude (one pooled SD), cost aggregated behavioral probe, frozen direction transfer to unseen feedback formats, and explicit persistence measurement. The central negative finding (decodability drops to 0.509 after eight unrelated tasks) is as valuable as the positive steering result, directly cautioning against treating any decodable valence direction as evidence of persistent welfare. The lexical control honestly quantifying that outcome wording contributes substantially to the direction is good scientific practice. Limitations are clearly stated: one model, one benchmark, non-monotone cost curves, feedback remains in history.
The idea of attempting to find a correlation between task success/failure and likelihood of wanting to reset the context is interesting. The main two issues I see with the current write-up are the potential for confounds in the chosen internal direction, and a lack of clarity in the methodology used. Given that MMLU is a knowledge-based multiple-choice test and fairly saturated as a benchmark, it's likely that the direction found is confounded with concepts such as correctness, honesty or faithfulness rather than "good task outcome"".
The "feedback wording semantic transfer" arm of the experiment could be presented more in detail; as it stands, it's difficult to understand what it entails.
As a general feedback, including prompt examples in the main test or an appendix would be helpful to clarify the experiment for readers.
Cite this work
@misc {
title={
(HckPrj) The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models
},
author={
Subramanyam Sahoo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


