Blue Beam
Patrick Dugan · Team MoralityLab
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
BlueBeam is a control-protocol research stack for risky software-agent behavior under imperfect monitorability. We built a benchmark-safe pipeline around BashArena-style trajectories that combines cheap risk heuristics, latent-sidechannel analysis, corruption-gap calibration, token-fed SAEs, and cross-referenced dictionary probes. The strongest current signal is orchestration-linked multi_step_risk_escalation, with useful supporting structure from manipulation-family probes such as unauthorized_config_edit and secret_touch_drift.
The positive result is that latent and probe-based analysis can rank behavior families and layers that separate benign from compromised conditions without saturating into hard-block behavior on benign traffic. The negative result is also included: a family-specific TinyLoRA follow-up did not yet translate into a strong controller-side intervention gain, so we present that as future work rather than overclaiming. We packaged figures, methodology, SAE outputs, and future- work notes in the submission bundle, and published a hosted artifact dashboard here: https://bluebeam-app.vercel.app

Reviews
- The calibrated controller idea is the strongest thread; build the paper around that.
- Focus the paper on one clear contribution with strong evidence rather than documenting an entire exploratory pipeline.
I think the main thing I would like is more numbers/metrics -- often I found myself wondering whether results/data were noise or actual signal. I think this was a rather ambitious project.
The other issue is that you're the one who tampers with the model in a specific way and then builds a detection system to find tampering and kind of that same axis. So it's unclear how well this would generalize and it does not seem to me as if it would generalize that well. You didn't talk much about generally baselining the original model.
I'd be interested in seeing some adversarial evaluations as well. It would be interesting to see a harmless fine tune and see if was not flagged.
Cite this project
@misc{dugan2026blue,
title = {{Blue Beam}},
author = {Patrick Dugan},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/blue-beam-yf5l}},
url = {https://apartresearch.com/sprints/projects/blue-beam-yf5l}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …