Skip to content
Sprint projectAug 16, 2026Cape Town

Persona or Model? Investigating Assistant Identity Through Persona-Conditioned Self-Reports

Kgothatso Mashigo , Themba Shongwe · Team Layered Minds

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Persona or Model? Investigating Assistant Identity Through Persona-Conditioned Self-Reports

Share

This project investigates whether persona instructions change how AI assistants describe themselves, or mainly affect the style of their responses. Gemini, ChatGPT, and Claude were each given analytical and creative-social personas and asked the same set of questions about identity, continuity, preferences, values, and autonomy. They were then asked to respond without considering the assigned persona. We compared their responses using qualitative coding to identify patterns and differences across models. While all three assistants generally separated persona from genuine personal traits and denied human-like continuity or subjective experience, they differed in how they described the relationship between persona and the underlying model. The project highlights both the usefulness and limitations of AI self-reports when studying model identity and AI welfare.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project adds to the field's black-box surface level examinations of personas within LLMs. Thoughtfully, this is a very readable paper which positions itself well philosophically yet lacks sufficient citations or positioning within the existing research. The methodology is clearly thoughtful and the focus on "claims about I, continuity, preference, and values" is well established. Models family names are noted as the subjects but their version numbers, dates, and inference context would be valuable details to include in future works. The paper claims a noted "pattern not seen described elsewhere" however several of the presentations during this sprint touched on and noted exactly this same type of pattern and often brought it into study at a greater depth, so the authors would be encouraged to attend and closely note the presentations during future sprints. Furthermore, Anthropic's work on The Persona Selection Model, and persona vectors work (Chen et al., 2025), would be valuable background to which contextualize this and start from a further place within the research. While this did not find a green field research problem to address, this is still a thoughtful contribution to the behavioral observations in the study of persona and identity in LLMs. Its explicit self-classification as an exploratory pilot and its refusal to treat self-reports as evidence of internals are practices worth emulating.

    Read full reviewShow less
  2. This is a clearly written exploratory study addressing an important methodological question for AI welfare and model-identity research: whether persona instructions alter substantive self-reports or mainly their stylistic presentation. The cross-model comparison is useful, and I particularly appreciated the authors' care in distinguishing observed self-reports from claims about underlying model architecture or consciousness.

    The main limitation is experimental replication. The most interesting finding—the difference between Gemini/ChatGPT's layered persona-over-model account and Claude's rejection of that framing—is currently based on one conversation per model/condition, so it is impossible to know whether this is a stable system-level difference or ordinary sampling/prompt variation. Repeating each condition across many fresh conversations, recording exact model versions, randomizing order, freezing identical prompts, and measuring the frequency with which each model produces each account would substantially strengthen the result.

    Independent qualitative coding is also important. A second blinded coder and inter-rater agreement would help validate the nine-pattern framework. Some comparisons should additionally be standardized: for example, the Claude repeated-probing behavior is difficult to interpret cross-model because the amount and wording of repetition differed between systems.

    A particularly informative follow-up would test whether each assistant can be induced to produce the other models' persona/model account under paraphrases and adversarial framing. If the layered-versus-non-layered distinction remains stable across independent samples and prompt variants, it would become a considerably stronger result. As presented, I view this as a useful pilot and protocol for generating hypotheses rather than strong evidence for stable cross-model identity differences.

    Read full reviewShow less

Cite this project

@misc{mashigo2026persona,
  title = {{Persona or Model? Investigating Assistant Identity Through Persona-Conditioned Self-Reports}},
  author = {Kgothatso Mashigo and Themba Shongwe},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/persona-or-model-investigating-assistant-identity-through-personaconditioned-selfreports-4lj7}},
  url = {https://apartresearch.com/sprints/projects/persona-or-model-investigating-assistant-identity-through-personaconditioned-selfreports-4lj7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026