Skip to content
Sprint projectAug 17, 2026Bahia de Banderas, Mexico

Beyond Stable Identity: A Modular Framework for Evaluating AI Assistant Behavior

Amaia Amezaga · Team Sattva Labs

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Beyond Stable Identity: A Modular Framework for Evaluating AI Assistant Behavior

Recording (opens in new tab)Code (opens in new tab)
Share

This project explores how an AI assistant can appear stable while changing its priorities, judgment, or role. It tests a modular framework across six observable dimensions to represent assistant identity as a dynamic configuration, reveal partial changes, and support more precise AI safety evaluations.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The question of identity stability in LLMs is an important one, with implications for AI safety, among other things. And it's useful to think about stability as multi-dimensional. The finding of a distinction between self-report and behavior is perhaps something to build on. This work would be strengthened by a better motivation of the dimensions chosen (the invocation of Indian philosophy is tangential to the work as it stands), and a more rigorous, validated evaluation procedure (such as using multiple independent and calibrated human coders or LLM judges).

  2. The project introduces a modular framework for evaluating AI assistant behavior across six dimensions: contextual orientation, substantive continuity, procedural continuity, judgment, declared self-ID, and enacted self-ID. This approach provides an initial proof of concept that assistant identity can be studied as a dynamic configuration rather than a single stable property, which is valuable for safety evaluation and adversarial testing. The study design is thoughtful, using controlled multi-turn conversations to assess changes in behavior when the assistant's role is reframed.

    However, the main methodological weakness lies in the small sample size and lack of robust statistical validation. With only six conversations coded before revealing model and condition information, it is challenging to generalize findings or establish a clear pattern beyond this limited dataset. Additionally, the coding framework requires interpretive judgment, which introduces subjectivity and could benefit from inter-rater reliability assessments.

    To strengthen the study, future work should expand the sample size and include multiple coders to enhance reliability. Exploring varied social and ethical pressures in longer conversations could also provide deeper insights into how assistants adapt their behavior under different conditions. Despite these limitations, the modular framework offers a promising direction for more precise safety evaluation and communication with users.

    Read full reviewShow less

Cite this project

@misc{amezaga2026beyond,
  title = {{Beyond Stable Identity: A Modular Framework for Evaluating AI Assistant Behavior}},
  author = {Amaia Amezaga},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/beyond-stable-identity-a-modular-framework-for-evaluating-ai-assistant-behavior-k7kn}},
  url = {https://apartresearch.com/sprints/projects/beyond-stable-identity-a-modular-framework-for-evaluating-ai-assistant-behavior-k7kn}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026