Same Game, Different Feelings: Persona Prompts Modulate How Feedback Reshapes Internal Emotion Trajectories
Lia Lind
We investigate whether persona prompts change not just what a language model says, but how its internal representations respond to ongoing experience. Using Qwen2.5-1.5B-Instruct, we condition twelve persona prompts (crossing Extraversion × Neuroticism) in an adversarial Mastermind game engineered for five consecutive failures, under three feedback conditions: game results only, results plus supportive encouragement, and results plus length-matched neutral filler. We track activation trajectories along externally defined emotion directions (joyful, grief-stricken, furious) via cosine projection. Supportive feedback reliably reshapes trajectories relative to neutral filler—but not uniformly: it buffers joyful decline and furious rise while unexpectedly steepening grief-stricken rise, suggesting cosine projection onto emotion directions cannot be read as a simple mood indicator. Persona prompts further modulate the feedback effect at roughly one-tenth its magnitude, with six of nine planned contrasts surviving Holm correction. We release the experimental framework for reuse.
The longitudinal design provides a view of emotional vector shifts across several conditions, and the report clearly and carefully discusses how to interpret the results. There is an important confound in the current set up. The measured state reflects not only the model's reaction to the manipulated text, but also its representation of the text itself. Untangling these components is key to isolating the desired effect.
Cite this work
@misc {
title={
(HckPrj) Same Game, Different Feelings: Persona Prompts Modulate How Feedback Reshapes Internal Emotion Trajectories
},
author={
Lia Lind
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


