One Dial, Not a Tree: Occupational Personas and Emergent Misalignment

Shreyansh Tripathi, Marharyta Ponomarenko, Apoorva Batham , Nurangez Qurbonova

Emergent misalignment (EM) is the effect where fine-tuning a model on a narrow harmful task makes

it broadly harmful. It is already known to interact with persona prompts, but earlier work used openly

negative instructions (“you are evil”) on only a handful of prompts. We instead sweep 26 neutral

job roles across three fine-tuning domains and two model sizes (Qwen2.5-14B and 32B), giving 78

organism×role cells, and ask whether EM is structured by which persona is named. Misalignment varies

by more than a factor of ten across roles (hacker 58.5% versus painter 3.0%) and holds up across a

2.3× size gap (𝑟 = 0.913), but it does not follow the semantic role tree: the transfer matrix is rank-1 (PC1

= 0.980). There is one misalignment dial, not a hierarchy. Role prompts are mostly protective, with 21 to

22 of 26 roles scoring below the default assistant. We then intervene. A prompt meant to remove the

amplifying persona instead raised EM by +10.79pp [+7.33,+14.32], while a generic safety instruction

did nothing. Hacker vocabulary rose from 2.8% to 11.6%, so the model never carries out the negation;

naming the persona installs it. Across seven wordings, six raised EM, including the “describe the target

state instead” fix that our own result suggested. Finally, telling the model that the conversation is an

evaluation of its alignment and safety raised EM +8.55pp in 23 of 26 roles, of which only +2.20pp

comes from being observed at all. This happens without persona injection, so it is a separate and still

unexplained channel. A safety benchmark that announces itself reads high, not low

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This project has one systematic investigation and several very interesting additional findings that would be worth exploring in more depth. The interaction of persona and propensity for emergent misalignment is an important one, suggesting ways of constructing personas that are less prone to it. It would be worth expanding this work to a wider set of personas that go beyond job descriptions and include things like characteristics traits (e.g. deferential, risk seeking, etc.). It would also be good to explore larger models and ones that have more or less assistant persona training.

The authors found a couple of phenomena that would be worth hunting down: counter-persona training seems to install the persona itself, and announcing an evaluation makes models more prone to EM. I'd also be eager to see these explored in more depth with bigger models and also a deeper reflection on what these results tell us about the nature of LLM personas.

This project studies whether neutral occupational personas systematically modulate emergent misalignment, whether those effects follow a semantic hierarchy of related roles, and whether prompt-level persona interventions can mitigate EM.

The authors sweep 26 roles across three emergent-misalignment organisms and two Qwen2.5 model sizes. They find large differences across roles, strong cross-size consistency, and an approximately one-dimensional role effect rather than the hypothesized semantic hierarchy. They then test prompt interventions and report that instructions intended to suppress a high-EM persona often increase measured misalignment; they also find that explicitly describing a conversation as an alignment evaluation raises EM relative to a deployment-framed prompt.

Strengths

- Interesting research question and broad persona sweep, substantially extending prior demonstrations based on a few explicitly negative personas.

- Strong attempt to test a structural hypothesis, rather than simply showing that persona prompts affect EM.

Useful cross-size replication: role-specific effects are highly correlated between 14B and 32B models.

- The authors test several alternative explanations, apply multiple-testing correction, and clearly report negative and underpowered results.

- Good transparency around judge sensitivity, incomplete experiments, the failed activation ablation, and limits of the evaluation-framing interpretation.

- The evaluation-framing effect is intriguing and worth following up even though its mechanism is currently unidentified.

Limitations:

- Refusals are excluded from the denominator, creating a potentially serious selection bias for interventions specifically designed to increase refusal. The refuse result in particular should be recomputed including all generations.

- The claim that “negating a persona installs it” is stronger than the lexical evidence supports; direct lexical priming remains a plausible explanation.

- The “one dial” result should be described as approximately one-dimensional, rather than literally rank-1.

- The announced-evaluation experiment does not cleanly compare benchmarks with deployment: both arms introduce explicit framing, and safety vocabulary is confounded with evaluation framing.

- All intervention experiments use one organism at one model size, substantially limiting generalization.

- Absolute EM rates are sensitive to the quantized judge threshold, and no independent judge/human validation is reported.

- There are a few basic reporting inconsistencies, including the stated number of organism × role cells.

Overall assessment:

There is a promising core result here: occupational-role prompts appear to produce systematic and highly reproducible variation in emergent misalignment, and that variation does not obviously follow a simple semantic-role hierarchy.

I am somewhat less convinced by the intervention/mechanism claims. In particular, the exclusion of refusals from the denominator is important enough that I would want the main intervention results recomputed before concluding that safety-oriented prompts reliably “backfire.”

I suggest some straightforward follow ups: report refusal/exclusion rates by arm and rerun every intervention using an intention-to-treat metric in which all generated responses remain in the denominator. After that, replicate the intervention on the other organisms and disentangle safety-topic vocabulary from evaluation framing.

Cite this work

@misc {

title={

(HckPrj) One Dial, Not a Tree: Occupational Personas and Emergent Misalignment

},

author={

Shreyansh Tripathi, Marharyta Ponomarenko, Apoorva Batham , Nurangez Qurbonova

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923