Shared Geometry, Causal Cross-Talk: Disentangling Persona and Emotion in Language Models
Varshith Vijjapu
We investigate whether persona and emotion representations in language models are independent or share causal structure. Across Qwen2.5-7B-Instruct and Granite-3.3-8B-Instruct, persona directions overlap substantially with emotion geometry, and across 20 emotions that overlap predicts downstream persona spillover during emotion steering. Removing only the persona-aligned component reduces spillover for 16/20 Qwen emotions and 17/20 Granite emotions while preserving 96.0% and 98.4% of the emotion effect. A controlled factorial experiment shows that the broader persona and valence representations nevertheless remain highly separable. Finally, a large causal perturbation of an internal valence direction barely changes a structured 0–9 emotional self-report, while persona and prompt framing strongly affect the report. These results highlight both mechanistic cross-talk and the need to causally validate model self-reports before treating them as evidence about internal affect.
This is a tightly-scoped, unusually rigorous mechanistic interpretability study. The core causal chain is well-constructed: persona directions occupy far more of the dominant emotion subspace than an isotropic baseline (16.4%/12.0% vs. 0.48%/0.59% expected), signed overlap predicts signed downstream persona spillover across a fixed 20-emotion set (r=.773/.841, positive in every one of 20 held-out evaluation questions), and norm-matched removal of only the persona-aligned component reduces that spillover for 16/20 and 17/20 emotions while retaining 96–98% of the intended emotion effect. The paper reports its own failures (specific emotions where orthogonalization didn't work) rather than only successes.
The factorial separability check (1,800 examples) rules out the obvious alternative explanation — that "persona" and "emotion" are just two noisy estimates of one representation — by showing both stay nearly invariant (cosine ~0.98–0.99) when the other factor is manipulated within the same story content.
The self-report finding is appropriately hedged as a calibration failure of one conversational assay, correctly positioned relative to the active 2026 introspection debate rather than overclaiming about subjective experience.
Generalization is the main limit, and the paper states this plainly: two model families at 7–8B scale, one operationalization of persona (sycophancy via five contrastive prompt pairs), linear interventions only, and causal measurement on fixed baseline tokens rather than free generation.
The pipeline replicates independently across two model families, interventions are norm matched, the factorial control is separate from the causal test, and the seven failures and the discarded behavioral judge are reported openly. The final-layer persistence of reduced spillover is the genuinely nontrivial part, since at the intervention layer the reduction holds by construction.
My concern is that the answer was largely predictable. A sycophancy direction should carry valence, the overlap signs mostly track valence polarity, and a locally linear propagation model already predicts both the correlation and the orthogonalization result. Spillover is also measured as projection onto the same vector p that defines the removed component, so intervention and readout share one geometry and the disentanglement result measures no behavior.
The most interesting data here are the failures. Qwen's hurt has overlap .001 and spillover of -4.8; Granite's thankful shows spillover of -7.8 on a +.022 overlap. Those are coupling channels this method cannot see, exactly what an auditor cares about, and they deserve analysis. Stage 6 is the most venue-relevant experiment and needs a positive control plus a real dose range before "calibration failure" means more than "insensitive at one dose magnitude."
The worthwhile next step is free generation on independently measured behaviors (honesty, refusal, deception, goal persistence) under adaptive prompting, or turning Stage 6 into a genuine calibration program for self-report assays. Either would make geometric decoupling load-bearing for alignment and digital minds, and the execution here says the author can do it.
Cite this work
@misc {
title={
(HckPrj) Shared Geometry, Causal Cross-Talk: Disentangling Persona and Emotion in Language Models
},
author={
Varshith Vijjapu
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


