The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models
Subramanyam Sahoo
This project studies whether language models develop internal representations of whether a task outcome went well or badly, how broadly those representations generalize, and whether they causally influence later choices. Using Qwen3 4B Instruct on MMLU, we identify a highly decodable outcome direction in the residual stream that transfers almost perfectly across several unseen ways of expressing success and failure. Although the representation is rapidly overwritten by unrelated context, activation steering along the same direction consistently shifts the model’s preference to reset its conversational state. Importantly, this behavioral effect appears without a detectable increase in global next token uncertainty. Together, the results reveal a compact and interpretable distinction between semantic generalization, contextual persistence, and causal influence. The project provides a simple experimental framework for studying functional outcome representations in language models while remaining agnostic about stronger claims concerning consciousness or phenomenal welfare.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models
},
author={
Subramanyam Sahoo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


