Beyond Stable Identity: A Modular Framework for Evaluating AI Assistant Behavior
Amaia Amezaga · Team Sattva Labs
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project explores how an AI assistant can appear stable while changing its priorities, judgment, or role. It tests a modular framework across six observable dimensions to represent assistant identity as a dynamic configuration, reveal partial changes, and support more precise AI safety evaluations.
Reviews
The question of identity stability in LLMs is an important one, with implications for AI safety, among other things. And it's useful to think about stability as multi-dimensional. The finding of a distinction between self-report and behavior is perhaps something to build on. This work would be strengthened by a better motivation of the dimensions chosen (the invocation of Indian philosophy is tangential to the work as it stands), and a more rigorous, validated evaluation procedure (such as using multiple independent and calibrated human coders or LLM judges).
The project introduces a modular framework for evaluating AI assistant behavior across six dimensions: contextual orientation, substantive continuity, procedural continuity, judgment, declared self-ID, and enacted self-ID. This approach provides an initial proof of concept that assistant identity can be studied as a dynamic configuration rather than a single stable property, which is valuable for safety evaluation and adversarial testing. The study design is thoughtful, using controlled multi-turn conversations to assess changes in behavior when the assistant's role is reframed.
However, the main methodological weakness lies in the small sample size and lack of robust statistical validation. With only six conversations coded before revealing model and condition information, it is challenging to generalize findings or establish a clear pattern beyond this limited dataset. Additionally, the coding framework requires interpretive judgment, which introduces subjectivity and could benefit from inter-rater reliability assessments.
To strengthen the study, future work should expand the sample size and include multiple coders to enhance reliability. Exploring varied social and ethical pressures in longer conversations could also provide deeper insights into how assistants adapt their behavior under different conditions. Despite these limitations, the modular framework offers a promising direction for more precise safety evaluation and communication with users.
Read full reviewShow less
Cite this project
@misc{amezaga2026beyond,
title = {{Beyond Stable Identity: A Modular Framework for Evaluating AI Assistant Behavior}},
author = {Amaia Amezaga},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/beyond-stable-identity-a-modular-framework-for-evaluating-ai-assistant-behavior-k7kn}},
url = {https://apartresearch.com/sprints/projects/beyond-stable-identity-a-modular-framework-for-evaluating-ai-assistant-behavior-k7kn}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …