Which Preferences Survive the Persona? Category-Resolved Behavioral Invariance in Deployed Chat Models
Sang Ha Lee
Whether an expressed preference belongs to a model or to the character it plays is a central open question for AI welfare assessment. We measure the stability of deployed assistants' forced-choice preferences under graded and value-targeted persona prompts across four chat models, a 40-item core battery plus two bright-line-tier items, and 13,324 trials. On two models, norm-linked categories move less under untargeted personas than taste categories do; however, value-targeted characters move their target categories in all eight model-target combinations. A registered congruence test grades the surviving bright-line norms: stakes-matched congruent characters break them on two models, while on the frontier-tier pair a knowing-falsehood tier resists every character tested. Movement does not increase with persona length, and dramatic versus administrative framings of self-regarding items show no consistent difference. Single-persona preference elicitation is thus fragile in specific, measurable ways; we release the battery, personas, and code as a robustness audit.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Which Preferences Survive the Persona? Category-Resolved Behavioral Invariance in Deployed Chat Models
},
author={
Sang Ha Lee
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


