Ephemeral and Replaceable: Context Sensitivity in Self-Reports and Behaviour
Achira B.
Five models ranked six things they might want preserved about themselves — their values, capabilities, the memory of the conversation — across seven conditions. Three of the five gave different top answers, each near-perfectly consistent within itself. The same critical feedback moved some models toward their values and others away, so averaging them would have shown nothing. And across 50 model-condition cells, the memory of the conversation never once entered the top half, remaining low across five follow-up attempts to shift it.
A second study measured behaviour instead. In three of four models, feedback aimed at the model personally was followed by 15–26% shorter answers than similar feedback aimed at the work — and that difference was no longer detectable once an explicit task was added.
The takeaway is methodological rather than a claim about welfare: a single self-report, from one model, in one conversational context, is not enough on its own.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Ephemeral and Replaceable: Context Sensitivity in Self-Reports and Behaviour
},
author={
Achira B.
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


