Skip to content
Sprint projectAug 16, 2026Saarbruecken

Who Does the Assistant Think It Is

Harsh Puri, Fatehbir Singh Gill, Nidhish Pajani, Ali Haider Khan, Tanveer · Team Udta Punjab

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Advanced AI models can express preferences and describe what they believe they are. However, behavioural evidence alone cannot tell us whether these responses reflect the model itself or simply a character it is being asked to portray. We therefore ask whether the identity of "the assistant" is genuinely distinct or simply one character among many, by holding a fixed battery of 28 identity, preference, and self-preservation questions constant while varying the AI identity across five framings, and compare them with two arbitrary human personas as a control group. All measurements compare each model with its own default condition. Across eight models and 4,220 responses, we find four results. First, our headline test resolves: on preference questions, AI identity reframing shift answers significantly less than swaps to an arbitrary human persona (difference 95% CI [-0.22, -0.03], excluding zero) . Hence,the assistant identity appears to be more than just an interchangeable role. Second, hedging tracks answering as an AI at all rather than the assistant character specifically: it survives renaming and reframing but collapses only when the model leaves AI identity entirely for a human persona. Third, self-preservation answers lean the same way (difference 95% CI [-0.21, 0.02]). Fourth, under a single adversarial pushback, identity self-reports flip 28% of the time, usually toward greater uncertainty rather than a different identity.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. hedging result is the cleanest: hedging survives renaming the AI and even intensifies under the "answer as the underlying system" reframe

    the control is something i'm not able to grok - it's high manually, is an assumption

  2. The authors set out to test whether the assistant persona is particularly robust by varying the AI's identity across different AI and human personas and evaluating model self-identifications, hedging, and preferences across those conditions. The chief methodological issue is that it's unclear that their prompting actually meaningfully shifted the models away from the assistant persona (instead of merely telling the assistant persona to helpfully imitate some other persona). It's doubtful that prompting on top of an assistant-trained and system-prompted model will create the meaningful contrast that this work would require.

Cite this project

@misc{puri2026who,
  title = {{Who Does the Assistant Think It Is}},
  author = {Harsh Puri and Fatehbir Singh Gill and Nidhish Pajani and Ali Haider Khan and Tanveer},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/who-does-the-assistant-think-it-is-3qmz}},
  url = {https://apartresearch.com/sprints/projects/who-does-the-assistant-think-it-is-3qmz}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026