Denoising Trajectories as Early-Warning Signals for Diffusion Language Model Control
Seunghyun Yoo · Team Carter
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
We study whether denoising trajectories in diffusion language models can serve as early warning signals for AI control. Using LLaDA with pretrained sparse autoencoders in an executable shell setting, we find trajectory-based monitoring detects attacks much earlier than output-only oversight, while steering shows weaker but preliminary promise for localized intervention.
Reviews
Using intermediary representations for monitoring in control and using steering is potentially interesting for control. Using diffusion LMs seems somewhat novel to me, however its difficult to see how this research is useful unless the most frontier models are also diffusion LMs.
This would have been stronger if run in a control environment like BigCodeBench or APPS (both are single turn). It's also hard to tell if the SAEs are actually classifying scheming or the input prompt. Also I didn't see many statistical tests / p values and I don't see why safety is unchanged using output only monitor vs no monitor. Both would have made this paper stronger.
Cite this project
@misc{yoo2026denoising,
title = {{Denoising Trajectories as Early-Warning Signals for Diffusion Language Model Control}},
author = {Seunghyun Yoo},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/denoising-trajectories-as-earlywarning-signals-for-diffusion-language-model-control-31cb}},
url = {https://apartresearch.com/sprints/projects/denoising-trajectories-as-earlywarning-signals-for-diffusion-language-model-control-31cb}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …