Who Does the Assistant Think It Is
Harsh Puri, Fatehbir Singh Gill, Nidhish Pajani, Ali Haider Khan, Tanveer
Advanced AI models can express preferences and describe what they believe they are. However, behavioural evidence alone cannot tell us whether these responses reflect the model itself or simply a character it is being asked to portray. We therefore ask whether the identity of "the assistant" is genuinely distinct or simply one character among many, by holding a fixed battery of 28 identity, preference, and self-preservation questions constant while varying the AI identity across five framings, and compare them with two arbitrary human personas as a control group. All measurements compare each model with its own default condition. Across eight models and 4,220 responses, we find four results. First, our headline test resolves: on preference questions, AI identity reframing shift answers significantly less than swaps to an arbitrary human persona (difference 95% CI [-0.22, -0.03], excluding zero) . Hence,the assistant identity appears to be more than just an interchangeable role. Second, hedging tracks answering as an AI at all rather than the assistant character specifically: it survives renaming and reframing but collapses only when the model leaves AI identity entirely for a human persona. Third, self-preservation answers lean the same way (difference 95% CI [-0.21, 0.02]). Fourth, under a single adversarial pushback, identity self-reports flip 28% of the time, usually toward greater uncertainty rather than a different identity.
The authors set out to test whether the assistant persona is particularly robust by varying the AI's identity across different AI and human personas and evaluating model self-identifications, hedging, and preferences across those conditions. The chief methodological issue is that it's unclear that their prompting actually meaningfully shifted the models away from the assistant persona (instead of merely telling the assistant persona to helpfully imitate some other persona). It's doubtful that prompting on top of an assistant-trained and system-prompted model will create the meaningful contrast that this work would require.
hedging result is the cleanest: hedging survives renaming the AI and even intensifies under the "answer as the underlying system" reframe
the control is something i'm not able to grok - it's high manually, is an assumption
Cite this work
@misc {
title={
(HckPrj) Who Does the Assistant Think It Is
},
author={
Harsh Puri, Fatehbir Singh Gill, Nidhish Pajani, Ali Haider Khan, Tanveer
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


