Track 1 - Vaultrix Control
XIAOHU ZHU · Team CSAGI
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
We presents five novel AI control evaluation settings designed to test control protocols against misaligned AI agents operating in realistic autonomous workloads in our real world Vaultrix project. The framework uses a tested Python library (vaultrix.control) compatible with ControlArena.
The design process revealed that the more autonomous an agent is, the more its legitimate actions overlap with attack patterns, creating an irreducible false-positive problem for monitors (the safety-usefulness tradeoff).
Multi-agent Difficulty: Multi-agent collusion (CollusionNet) is qualitatively harder to detect than single-agent attacks because the attack is distributed across agents and embedded in shared artifacts, a gap current protocols do not support.
Persistence Risk: Persistence mechanisms (systemd services, cron jobs) pose the greatest unaddressed risk, as they create harm that outlasts the agent's session and bypass typical post-use container teardowns.
The current limitations are that the monitors are only regex-based baselines and would be bypassed by a sophisticated agent that obfuscates its actions. The tasks are currently evaluated via artifact dictionaries rather than live execution in fully dockerized environments.
Reviews
Five threat model designs, CollusionNet and RogueInterp especially are creative and fill real gaps. But no models were run against any of them. No attack rates, no detection rates, nothing empirical. Code works with 67 passing tests but that's not results. One validated setting with actual data would be worth more than five designs on paper.
Cite this project
@misc{zhu2026track,
title = {{Track 1 - Vaultrix Control}},
author = {XIAOHU ZHU},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/track-1-vaultrix-control-twqz}},
url = {https://apartresearch.com/sprints/projects/track-1-vaultrix-control-twqz}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …