Katechon: Empirical Benchmarking of AI Control Protocols in a Proactive Defense Setting
Volkan Erdogan
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
An open-source framework for empirically benchmarking AI control protocols. A closed OODA loop generates threat intervention plans while an adversarial planner injects subtly harmful alternatives, reframing the planner as an untrusted model. Four pluggable control protocols are evaluated across eight reproducible scenarios. Key findings: threshold-based gating is blind to adversarial content (TPR 0%), constitutional constraints achieve the best safety-usefulness balance (TPR 46.67%, FPR 5.71%), and debate maximizes detection (TPR 100%) at steep false-positive cost (FPR 44.64%). Stealth adversarial plans evade all protocols, identifying longitudinal anomaly detection as a key open problem.
Reviews
Nice implementation of a multi-turn policy, and nice approach.
One minor engineering nit: this is poorly documented, but Luthien policy's are (supposed to be) ~stateless, as the same policy can be acting on many requests at once and you usually don't want those results interfering with each other - there are ways to persist state across turns that are session-specific, but they're poorly documented. This is mostly a criticism of Luthien's documentation at the moment, rather than your implementation.
Cite this project
@misc{erdogan2026katechon,
title = {{Katechon: Empirical Benchmarking of AI Control Protocols in a Proactive Defense Setting}},
author = {Volkan Erdogan},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/katechon-empirical-benchmarking-of-ai-control-protocols-in-a-proactive-defense-setting-kxbo}},
url = {https://apartresearch.com/sprints/projects/katechon-empirical-benchmarking-of-ai-control-protocols-in-a-proactive-defense-setting-kxbo}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …