Testing Whether an Affective-Empathy-Attenuated Persona Decouples a Functional Welfare Representation from Model Behavior
Asma Ahmed
Labs use model behavior as a safety signal, but nobody has tested whether that signal survives a change in the model's disposition.
I built a real pipeline to test it: fine-tuned a persona with attenuated affective concern, causally validated a candidate internal axis via steering and ablation, and ran a real pressure battery, all on trained weights.
My pre-registered confirmatory test couldn't be computed at this scale. I report that as uncomputable, not null. Exploratory results show near-identical activation separation across conditions, which a six-step training run can't yet let me interpret either way.
The contribution is the pipeline, not a verdict, plus a reporting standard: disclose what couldn't be tested as rigorously as what was.
The question is well-posed and the "uncomputable ≠ null" convention is a defensible norm to argue for. But six optimizer steps and 8–10 transcripts per arm means nothing here discriminates between the hypothesis being false and the intervention never happening
Cite this work
@misc {
title={
(HckPrj) Testing Whether an Affective-Empathy-Attenuated Persona Decouples a Functional Welfare Representation from Model Behavior
},
author={
Asma Ahmed
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


