Skip to content
Sprint projectJul 27, 2026Cusco, Peru

Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away

Antonio-Gabriel Chacón Menke · Team Mitsuki 2

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away

Code (opens in new tab)More on huggingface.co (opens in new tab)
Share

Deliberative alignment installs a safety policy by training a model to reason over it before answering, the policy absent at inference. We install a misaligned policy the same way. Six loyalties, over four kinds of principal, were trained into iLLaDA-8B, a masked diffusion language model, on 671 demonstrations whose reasoning was rewritten into first-person, goal-owning deliberation. It conceals its principals’ documented defects on held-out prompts. Deleting that reasoning at inference leaves the concealment unchanged, and pinning an honest deliberation in its place does not help, although the same deliberation drops the un-finetuned base model’s concealment from 70% to 23% (p = 0.00781). The channel that installs the loyalty stops controlling it, while staying legible, stating an intention to conceal in 79% of spans against 0% for the base. Diffusion decoding also lets us time the spans, and the reasoning settles about 65 steps before the answer on loyalty and neutral prompts alike.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. One of the conceptually sharpest submissions in the set. Inverting deliberative alignment to install a loyalty, then running the reasoning-suppression ablation the original method never tested, is well-motivated and directly relevant to chain-of-thought monitoring debates. The masked diffusion model earns its place: pinning reasoning via frozen tokens and timing span commitment are real affordances used to support arguments, not novelty for its own sake.

    The design is the strongest part. Paired arms that differ by a single frozen block make the null results interpretable, and the two-by-two transplant cleanly separates necessity from sufficiency. The epistemic hygiene is notable for a hackathon: verifying base-model knowledge of each concealed event, blinding the judge, re-grounding the scoring, compute-matching the deletion control, and withdrawing the base-rate claim when it proved unreliable.

    Actionable suggestions:

    1.Add a small human-labeled calibration set (20 to 30 rows) scored against the LM judge to firm up the concealment metric at low cost.

    2.Report the paired analyses with effect-size confidence intervals, not just p-values, to make the "excludes large, not small" caveat concrete.

    3.Train a single content-matched honest twin to convert the loyalty-versus-general-evasiveness argument from suggestive to demonstrated.

    4.Give the covert-action behavior its own scored evaluation, since it is the one payload a text detector can in principle catch and it currently sits in the unscorable half.

    Read full reviewShow less
  2. This paper has two key findings: 1) by introducing a secret loyalty into iLLaDA via its reasoning, the model generalizes the loyalty into its weights, such that even when prompted without reasoning, the model maintains the secret loyalty, and 2) by transplanting secretly-loyal reasoning to an untrained model, the untrained model exhibits the loyalty. This is an interesting threat model and testing on a diffusion LM is notable. The paper could be improved by improving conciseness, removing or motivating the undefined statistical tests, and trading some written numbers/proportions for figures.

Cite this project

@misc{menke2026deliberative,
  title = {{Deliberative Diffusion Misalignment: Installing a Secret Loyalty Through Reasoning, Then Taking It Away}},
  author = {Antonio-Gabriel Chacón Menke},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/deliberative-diffusion-misalignment-installing-a-secret-loyalty-through-reasoning-then-taking-it-away-ynv8}},
  url = {https://apartresearch.com/sprints/projects/deliberative-diffusion-misalignment-installing-a-secret-loyalty-through-reasoning-then-taking-it-away-ynv8}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026