Shared Geometry, Causal Cross-Talk: Disentangling Persona and Emotion in Language Models

Varshith Vijjapu

We investigate whether persona and emotion representations in language models are independent or share causal structure. Across Qwen2.5-7B-Instruct and Granite-3.3-8B-Instruct, persona directions overlap substantially with emotion geometry, and across 20 emotions that overlap predicts downstream persona spillover during emotion steering. Removing only the persona-aligned component reduces spillover for 16/20 Qwen emotions and 17/20 Granite emotions while preserving 96.0% and 98.4% of the emotion effect. A controlled factorial experiment shows that the broader persona and valence representations nevertheless remain highly separable. Finally, a large causal perturbation of an internal valence direction barely changes a structured 0–9 emotional self-report, while persona and prompt framing strongly affect the report. These results highlight both mechanistic cross-talk and the need to causally validate model self-reports before treating them as evidence about internal affect.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is a tightly-scoped, unusually rigorous mechanistic interpretability study. The core causal chain is well-constructed: persona directions occupy far more of the dominant emotion subspace than an isotropic baseline (16.4%/12.0% vs. 0.48%/0.59% expected), signed overlap predicts signed downstream persona spillover across a fixed 20-emotion set (r=.773/.841, positive in every one of 20 held-out evaluation questions), and norm-matched removal of only the persona-aligned component reduces that spillover for 16/20 and 17/20 emotions while retaining 96–98% of the intended emotion effect. The paper reports its own failures (specific emotions where orthogonalization didn't work) rather than only successes.

The factorial separability check (1,800 examples) rules out the obvious alternative explanation — that "persona" and "emotion" are just two noisy estimates of one representation — by showing both stay nearly invariant (cosine ~0.98–0.99) when the other factor is manipulated within the same story content.

The self-report finding is appropriately hedged as a calibration failure of one conversational assay, correctly positioned relative to the active 2026 introspection debate rather than overclaiming about subjective experience.

Generalization is the main limit, and the paper states this plainly: two model families at 7–8B scale, one operationalization of persona (sycophancy via five contrastive prompt pairs), linear interventions only, and causal measurement on fixed baseline tokens rather than free generation.

The pipeline replicates independently across two model families, interventions are norm matched, the factorial control is separate from the causal test, and the seven failures and the discarded behavioral judge are reported openly. The final-layer persistence of reduced spillover is the genuinely nontrivial part, since at the intervention layer the reduction holds by construction.

My concern is that the answer was largely predictable. A sycophancy direction should carry valence, the overlap signs mostly track valence polarity, and a locally linear propagation model already predicts both the correlation and the orthogonalization result. Spillover is also measured as projection onto the same vector p that defines the removed component, so intervention and readout share one geometry and the disentanglement result measures no behavior.

The most interesting data here are the failures. Qwen's hurt has overlap .001 and spillover of -4.8; Granite's thankful shows spillover of -7.8 on a +.022 overlap. Those are coupling channels this method cannot see, exactly what an auditor cares about, and they deserve analysis. Stage 6 is the most venue-relevant experiment and needs a positive control plus a real dose range before "calibration failure" means more than "insensitive at one dose magnitude."

The worthwhile next step is free generation on independently measured behaviors (honesty, refusal, deception, goal persistence) under adaptive prompting, or turning Stage 6 into a genuine calibration program for self-report assays. Either would make geometric decoupling load-bearing for alignment and digital minds, and the execution here says the author can do it.

Cite this work

@misc {

title={

(HckPrj) Shared Geometry, Causal Cross-Talk: Disentangling Persona and Emotion in Language Models

},

author={

Varshith Vijjapu

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923