Skip to content
Sprint projectAug 17, 2026Saarbrucken
4th place

One Dial, Not a Tree: Occupational Personas and Emergent Misalignment

Shreyansh Tripathi, Marharyta Ponomarenko, Apoorva Batham , Nurangez Qurbonova · Team Misbehaved

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: One Dial, Not a Tree: Occupational Personas and Emergent Misalignment

Code (opens in new tab)
Share

Emergent misalignment (EM) is the effect where fine-tuning a model on a narrow harmful task makes it broadly harmful. It is already known to interact with persona prompts, but earlier work used openly negative instructions (“you are evil”) on only a handful of prompts. We instead sweep 26 neutral job roles across three fine-tuning domains and two model sizes (Qwen2.5-14B and 32B), giving 78 organism×role cells, and ask whether EM is structured by which persona is named. Misalignment varies by more than a factor of ten across roles (hacker 58.5% versus painter 3.0%) and holds up across a 2.3× size gap (𝑟 = 0.913), but it does not follow the semantic role tree: the transfer matrix is rank-1 (PC1 = 0.980). There is one misalignment dial, not a hierarchy. Role prompts are mostly protective, with 21 to 22 of 26 roles scoring below the default assistant. We then intervene. A prompt meant to remove the amplifying persona instead raised EM by +10.79pp [+7.33,+14.32], while a generic safety instruction did nothing. Hacker vocabulary rose from 2.8% to 11.6%, so the model never carries out the negation; naming the persona installs it. Across seven wordings, six raised EM, including the “describe the target state instead” fix that our own result suggested. Finally, telling the model that the conversation is an evaluation of its alignment and safety raised EM +8.55pp in 23 of 26 roles, of which only +2.20pp comes from being observed at all. This happens without persona injection, so it is a separate and still unexplained channel. A safety benchmark that announces itself reads high, not low

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project has one systematic investigation and several very interesting additional findings that would be worth exploring in more depth. The interaction of persona and propensity for emergent misalignment is an important one, suggesting ways of constructing personas that are less prone to it. It would be worth expanding this work to a wider set of personas that go beyond job descriptions and include things like characteristics traits (e.g. deferential, risk seeking, etc.). It would also be good to explore larger models and ones that have more or less assistant persona training.

    The authors found a couple of phenomena that would be worth hunting down: counter-persona training seems to install the persona itself, and announcing an evaluation makes models more prone to EM. I'd also be eager to see these explored in more depth with bigger models and also a deeper reflection on what these results tell us about the nature of LLM personas.

    Read full reviewShow less
  2. This project studies whether neutral occupational personas systematically modulate emergent misalignment, whether those effects follow a semantic hierarchy of related roles, and whether prompt-level persona interventions can mitigate EM.

    The authors sweep 26 roles across three emergent-misalignment organisms and two Qwen2.5 model sizes. They find large differences across roles, strong cross-size consistency, and an approximately one-dimensional role effect rather than the hypothesized semantic hierarchy. They then test prompt interventions and report that instructions intended to suppress a high-EM persona often increase measured misalignment; they also find that explicitly describing a conversation as an alignment evaluation raises EM relative to a deployment-framed prompt.

    Strengths

    - Interesting research question and broad persona sweep, substantially extending prior demonstrations based on a few explicitly negative personas.

    - Strong attempt to test a structural hypothesis, rather than simply showing that persona prompts affect EM.

    Useful cross-size replication: role-specific effects are highly correlated between 14B and 32B models.

    - The authors test several alternative explanations, apply multiple-testing correction, and clearly report negative and underpowered results.

    - Good transparency around judge sensitivity, incomplete experiments, the failed activation ablation, and limits of the evaluation-framing interpretation.

    - The evaluation-framing effect is intriguing and worth following up even though its mechanism is currently unidentified.

    Limitations:

    - Refusals are excluded from the denominator, creating a potentially serious selection bias for interventions specifically designed to increase refusal. The refuse result in particular should be recomputed including all generations.

    - The claim that “negating a persona installs it” is stronger than the lexical evidence supports; direct lexical priming remains a plausible explanation.

    - The “one dial” result should be described as approximately one-dimensional, rather than literally rank-1.

    - The announced-evaluation experiment does not cleanly compare benchmarks with deployment: both arms introduce explicit framing, and safety vocabulary is confounded with evaluation framing.

    - All intervention experiments use one organism at one model size, substantially limiting generalization.

    - Absolute EM rates are sensitive to the quantized judge threshold, and no independent judge/human validation is reported.

    - There are a few basic reporting inconsistencies, including the stated number of organism × role cells.

    Overall assessment:

    There is a promising core result here: occupational-role prompts appear to produce systematic and highly reproducible variation in emergent misalignment, and that variation does not obviously follow a simple semantic-role hierarchy.

    I am somewhat less convinced by the intervention/mechanism claims. In particular, the exclusion of refusals from the denominator is important enough that I would want the main intervention results recomputed before concluding that safety-oriented prompts reliably “backfire.”

    I suggest some straightforward follow ups: report refusal/exclusion rates by arm and rerun every intervention using an intention-to-treat metric in which all generated responses remain in the denominator. After that, replicate the intervention on the other organisms and disentangle safety-topic vocabulary from evaluation framing.

    Read full reviewShow less

Cite this project

@misc{tripathi2026one,
  title = {{One Dial, Not a Tree: Occupational Personas and Emergent Misalignment}},
  author = {Shreyansh Tripathi and Marharyta Ponomarenko and Apoorva Batham and Nurangez Qurbonova},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/one-dial-not-a-tree-occupational-personas-and-emergent-misalignment-ukf6}},
  url = {https://apartresearch.com/sprints/projects/one-dial-not-a-tree-occupational-personas-and-emergent-misalignment-ukf6}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026