Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away
Antonio-Gabriel Chacón Menke · Team Mitsuki 2
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Deliberative alignment installs a safety policy by training a model to reason over it before answering, the policy absent at inference. We install a misaligned policy the same way. Six loyalties, over four kinds of principal, were trained into iLLaDA-8B, a masked diffusion language model, on 671 demonstrations whose reasoning was rewritten into first-person, goal-owning deliberation. It conceals its principals’ documented defects on held-out prompts. Deleting that reasoning at inference leaves the concealment unchanged, and pinning an honest deliberation in its place does not help, although the same deliberation drops the un-finetuned base model’s concealment from 70% to 23% (p = 0.00781). The channel that installs the loyalty stops controlling it, while staying legible, stating an intention to conceal in 79% of spans against 0% for the base. Diffusion decoding also lets us time the spans, and the reasoning settles about 65 steps before the answer on loyalty and neutral prompts alike.
Reviews
One of the conceptually sharpest submissions in the set. Inverting deliberative alignment to install a loyalty, then running the reasoning-suppression ablation the original method never tested, is well-motivated and directly relevant to chain-of-thought monitoring debates. The masked diffusion model earns its place: pinning reasoning via frozen tokens and timing span commitment are real affordances used to support arguments, not novelty for its own sake.
The design is the strongest part. Paired arms that differ by a single frozen block make the null results interpretable, and the two-by-two transplant cleanly separates necessity from sufficiency. The epistemic hygiene is notable for a hackathon: verifying base-model knowledge of each concealed event, blinding the judge, re-grounding the scoring, compute-matching the deletion control, and withdrawing the base-rate claim when it proved unreliable.
Actionable suggestions:
1.Add a small human-labeled calibration set (20 to 30 rows) scored against the LM judge to firm up the concealment metric at low cost.
2.Report the paired analyses with effect-size confidence intervals, not just p-values, to make the "excludes large, not small" caveat concrete.
3.Train a single content-matched honest twin to convert the loyalty-versus-general-evasiveness argument from suggestive to demonstrated.
4.Give the covert-action behavior its own scored evaluation, since it is the one payload a text detector can in principle catch and it currently sits in the unscorable half.
Read full reviewShow less
This paper has two key findings: 1) by introducing a secret loyalty into iLLaDA via its reasoning, the model generalizes the loyalty into its weights, such that even when prompted without reasoning, the model maintains the secret loyalty, and 2) by transplanting secretly-loyal reasoning to an untrained model, the untrained model exhibits the loyalty. This is an interesting threat model and testing on a diffusion LM is notable. The paper could be improved by improving conciseness, removing or motivating the undefined statistical tests, and trading some written numbers/proportions for figures.
Cite this project
@misc{menke2026deliberative,
title = {{Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away}},
author = {Antonio-Gabriel Chacón Menke},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/deliberative-diffusion-misalignment-installing-a-secret-loyalty-through-reasoning-then-taking-it-away-ynv8}},
url = {https://apartresearch.com/sprints/projects/deliberative-diffusion-misalignment-installing-a-secret-loyalty-through-reasoning-then-taking-it-away-ynv8}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …