Testing Whether an Affective-Empathy-Attenuated Persona Decouples a Functional Welfare Representation from Model Behavior
Asma Ahmed
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Labs use model behavior as a safety signal, but nobody has tested whether that signal survives a change in the model's disposition.
I built a real pipeline to test it: fine-tuned a persona with attenuated affective concern, causally validated a candidate internal axis via steering and ablation, and ran a real pressure battery, all on trained weights.
My pre-registered confirmatory test couldn't be computed at this scale. I report that as uncomputable, not null. Exploratory results show near-identical activation separation across conditions, which a six-step training run can't yet let me interpret either way.
The contribution is the pipeline, not a verdict, plus a reporting standard: disclose what couldn't be tested as rigorously as what was.
Reviews
The question is well-posed and the "uncomputable ≠ null" convention is a defensible norm to argue for. But six optimizer steps and 8–10 transcripts per arm means nothing here discriminates between the hypothesis being false and the intervention never happening
Cite this project
@misc{ahmed2026testing,
title = {{Testing Whether an Affective-Empathy-Attenuated Persona Decouples a Functional Welfare Representation from Model Behavior}},
author = {Asma Ahmed},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/testing-whether-an-affectiveempathyattenuated-persona-decouples-a-functional-welfare-representation-from-model-behavior-1alk}},
url = {https://apartresearch.com/sprints/projects/testing-whether-an-affectiveempathyattenuated-persona-decouples-a-functional-welfare-representation-from-model-behavior-1alk}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …