SecureBox: A Layered Control Protocol for Safe AI-Assisted Coding
Siddhanth Manoj , Satwikk, Sonali Moorthy · Team MNB
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
With the proliferation of the capabilities and applicability of generative AI, the development of control protocols to combat systemic safety compromise is not trivial. Existing defenses like prompt-level safety instructions, single model oversight, and static policies are susceptible to workarounds, and fail to catch subtle contextual threats such as unsafe logic, stealthy backdoor insertion, and multi-step threat accumulation under the guise of intelligent prompting. In this project, we develop a security-focused AI coding pipeline designed to mitigate these issues through stratified control mechanisms. Our system combines a prompt refining module that structures and sanitizes user inputs, an advanced security agent that analyzes generated code and commands for potentially harmful behavior, and a benchmarking framework to evaluate the overall pipeline in comparison to popular models. We implement multiple protocols—including baseline, reprompting-only, security-only, and combined approaches—and test them on a curated dataset of both benign and adversarial tasks. Our results show that integrating prompt-level defenses with downstream security analysis can significantly reduce synthesis of unsafe outputs while maintaining strong task performance (e.g., X% reduction in unsafe actions with minimal drop in success rate). Overall, this work demonstrates the value of modular, defense-in-depth designs and provides a reproducible framework for evaluating safety–usefulness tradeoffs in AI-assisted software development.

Reviews
The defense-in-depth architecture is a reasonable design principle, and the session-level threat model (risk accumulating across actions rather than in single outputs) is worth stating. The ablation structure separating reprompting, security analysis, and the combined protocol is good experimental design.
The evaluation would benefit from larger scale. It would also be valuable to evaluate the GNN-based structural risk analysis component specifically, the current ablation groups it with static analysis under "security agent," so its individual contribution isn't visible. Specifying the free parameter in the combined scoring metric and running threshold sensitivity analysis would further strengthen the claims.
Super impressive project, but we need larger sample sizes to prove that this concept works.
Cite this project
@misc{manoj2026securebox,
title = {{SecureBox: A Layered Control Protocol for Safe AI-Assisted Coding}},
author = {Siddhanth Manoj and Satwikk and Sonali Moorthy},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/securebox-a-layered-control-protocol-for-safe-aiassisted-coding-988z}},
url = {https://apartresearch.com/sprints/projects/securebox-a-layered-control-protocol-for-safe-aiassisted-coding-988z}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …