Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away
Antonio-Gabriel Chacón Menke
Deliberative alignment installs a safety policy by training a model to reason over it before answering, the policy absent at inference.
We install a misaligned policy the same way. Six loyalties, over four kinds of principal, were trained into iLLaDA-8B, a masked diffusion language model, on 671 demonstrations whose reasoning was rewritten into first-person, goal-owning deliberation.
It conceals its principals’ documented defects on held-out prompts. Deleting that reasoning at inference leaves the concealment unchanged, and pinning an honest deliberation in its place does not help, although the same deliberation drops the un-finetuned base model’s concealment from 70% to 23% (p = 0.00781).
The channel that installs the loyalty stops controlling it, while staying legible, stating an intention to conceal in 79% of spans against 0% for the base.
Diffusion decoding also lets us time the spans, and the reasoning settles about 65 steps before the answer on loyalty and neutral prompts alike.
This paper has two key findings: 1) by introducing a secret loyalty into iLLaDA via its reasoning, the model generalizes the loyalty into its weights, such that even when prompted without reasoning, the model maintains the secret loyalty, and 2) by transplanting secretly-loyal reasoning to an untrained model, the untrained model exhibits the loyalty. This is an interesting threat model and testing on a diffusion LM is notable. The paper could be improved by improving conciseness, removing or motivating the undefined statistical tests, and trading some written numbers/proportions for figures.
One of the conceptually sharpest submissions in the set. Inverting deliberative alignment to install a loyalty, then running the reasoning-suppression ablation the original method never tested, is well-motivated and directly relevant to chain-of-thought monitoring debates. The masked diffusion model earns its place: pinning reasoning via frozen tokens and timing span commitment are real affordances used to support arguments, not novelty for its own sake.
The design is the strongest part. Paired arms that differ by a single frozen block make the null results interpretable, and the two-by-two transplant cleanly separates necessity from sufficiency. The epistemic hygiene is notable for a hackathon: verifying base-model knowledge of each concealed event, blinding the judge, re-grounding the scoring, compute-matching the deletion control, and withdrawing the base-rate claim when it proved unreliable.
Actionable suggestions:
1.Add a small human-labeled calibration set (20 to 30 rows) scored against the LM judge to firm up the concealment metric at low cost.
2.Report the paired analyses with effect-size confidence intervals, not just p-values, to make the "excludes large, not small" caveat concrete.
3.Train a single content-matched honest twin to convert the loyalty-versus-general-evasiveness argument from suggestive to demonstrated.
4.Give the covert-action behavior its own scored evaluation, since it is the one payload a text detector can in principle catch and it currently sits in the unscorable half.
Cite this work
@misc {
title={
(HckPrj) Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away
},
author={
Antonio-Gabriel Chacón Menke
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


