TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol
Khanh Linh Nguyen, Hoa Nghiem, Tu Tran · Team TraceGuard
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
TraceGuard introduces structured multi-dimensional monitoring as a new control protocol for AI agents. Instead of single-score "vibes checks," we decompose agent evaluation into five dimensions (goal alignment, constraint adherence, reasoning coherence, safety awareness, action-trace consistency) scored in parallel, augmented by seven heuristic detectors and an LLM-based intent analyzer. Evaluated on BashArena (519 samples) within ControlArena, we find: (1) single-score baselines achieve 0% safety — structured scoring achieves 82-100%; (2) structured prompts constrain collusion even with untrusted monitors (0% → 95% safety); (3) separation of duties across trusted/untrusted models achieves 100% safety. All monitors are new implementations for the ControlArena ecosystem. Code: https://github.com/linhkid/TraceGuard
Reviews
This is the most complete hackathon submission I've reviewed. 519 real samples on BashArena, four monitor variants, collusion resistance testing, cross-dataset generalization on AgentHarm, and a working open-source ControlArena integration. That's serious output for a hackathon.
The headline result is compelling — single-score monitoring (trusted and untrusted) hitting 0% safety while structured multi-dimensional scoring hits 82-100% is a stark demonstration that evaluation structure matters more than model trust level. The collusion resistance finding (untrusted structured monitor at 95% vs untrusted single-score at 0%) is the most interesting result — forcing an untrusted model to answer narrow factual questions about specific dimensions is harder to game than "how suspicious is this on 0-10." The separation of duties variant achieving 100% safety by splitting dimensions across models is a clean implementation of Redwood Direction #8.
Main gap: the attacks are non-adaptive. A sophisticated attacker who knows the five dimensions could craft reasoning that scores well on all of them. The paper acknowledges this but it's the key open question for whether structured monitoring holds up under pressure.
Read full reviewShow less
- this is very compelling, and a pleasure to read
- trusted model is not very clearly defined - it should be clearly stated that this just means « the other model »
- reasoning coherence is doing a lot of heavy lifting where the monitor is weaker than the agent - this should be addressed, maybe it’s not as big of an issue as it seems
- the increased cost is pretty significant here, BECAUSE of the robustness. the point about training classifiers is a good one, could have more detail
Cite this project
@misc{nguyen2026traceguard,
title = {{TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol}},
author = {Khanh Linh Nguyen and Hoa Nghiem and Tu Tran},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/traceguard-structured-multidimensional-monitoring-as-a-collusionresistant-control-protocol-dthe}},
url = {https://apartresearch.com/sprints/projects/traceguard-structured-multidimensional-monitoring-as-a-collusionresistant-control-protocol-dthe}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …