Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away
Antonio-Gabriel Chacón Menke
Deliberative alignment installs a safety policy by training a model to reason over it before answering, the policy absent at inference.
We install a misaligned policy the same way. Six loyalties, over four kinds of principal, were trained into iLLaDA-8B, a masked diffusion language model, on 671 demonstrations whose reasoning was rewritten into first-person, goal-owning deliberation.
It conceals its principals’ documented defects on held-out prompts. Deleting that reasoning at inference leaves the concealment unchanged, and pinning an honest deliberation in its place does not help, although the same deliberation drops the un-finetuned base model’s concealment from 70% to 23% (p = 0.00781).
The channel that installs the loyalty stops controlling it, while staying legible, stating an intention to conceal in 79% of spans against 0% for the base.
Diffusion decoding also lets us time the spans, and the reasoning settles about 65 steps before the answer on loyalty and neutral prompts alike.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away
},
author={
Antonio-Gabriel Chacón Menke
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


