Constraint Dilemma
Manu Xaviour Thaisseril Shaju, Soham Takawadekar · Team Manu- Soham
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
Trusted monitors can catch backdoors in AI-generated code, but only by flagging a large share of honest submissions for costly human review. We propose enforced comment constraints — mandatory, structured documentation rules that create a dilemma for adversarial models: follow the constraints, which bind comments to actual code semantics, exposing the backdoor to the monitor — or break them and get caught by an automated compliance checker before the monitor even runs. Building on Sreevatsa et al. (2026), who showed that optional comments improve trusted monitoring, we make comments mandatory and structured, eliminating the attacker's ability to write vague or misleading documentation. We introduce a 5-class constraint taxonomy ranging from free AST-based presence checks to LLM-verified semantic accuracy, and evaluate it against six adversarial commenting strategies across 200 coding problems. Our experimental data is synthetic due to compute limitations; the contribution is the constraint dilemma framework and methodology, not the specific numbers
Reviews
The constraint dilemma framing is clever — forcing adversarial models to either honestly describe their backdoor (exposing it to the monitor) or violate the format and get caught by an automated checker before the monitor even runs. No third option. The 5-class taxonomy is well-structured, and the observation that Classes 1-2 (AST parsing, regex) are free and deterministic makes the cheapest version immediately deployable. Even just requiring branch-level comments eliminates the entire "leave the backdoor uncommented" strategy class at zero cost.
The main issue is that the entire evaluation is synthetic — no real model outputs, no real constraint checking, no real monitoring. The authors are upfront about this (GPU availability), which I respect, but it means the heatmaps and coverage matrices are LLM-generated simulations of what results might look like, not evidence that the dilemma actually works. The backdoor model tested (if n == 47: return -1) is also very simple — real backdoors look like honest mistakes, not obvious triggers, and it's unclear whether the constraint taxonomy holds against adversaries who write comments that are literally true but selectively emphasize the wrong things.
Good framework, solid writeup, needs real validation on ControlArena to land.
Read full reviewShow less
Nice framing for enforcing comment standards to support control protocols and monitoring. Due to compute constraints the current results are unfortunately simulated and uninformative, but the approach would be interesting to explore in practice. In particular measuring the increase in FPR from the constraints, and the ability of a powerful model to produce adversarial comments that both pass constraints checks and fool the monitor.
Cite this project
@misc{shaju2026constraint,
title = {{Constraint Dilemma}},
author = {Manu Xaviour Thaisseril Shaju and Soham Takawadekar},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/constraint-dilemma-udka}},
url = {https://apartresearch.com/sprints/projects/constraint-dilemma-udka}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …