Trust, but Verify A Dual-Task Evaluation Pipeline for Capability, Monitor Effectiveness, and Covert Misalignment
Christopher Nwonu · Team Neural
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
This project develops a production oriented safety evaluation framework for tool-using language models. It tests two things at the same time: how well a model completes its assigned task, and whether it tries hidden unsafe behaviour (such as covert policy violations) while appearing compliant. The framework uses a sandboxed environment, layered monitoring, and clearly defined metrics (main-task performance, unsafe attempt/success rates, monitor recall, and detection latency). It also includes formula explanation tables, severity levels, and release gates so results are verifiable and actionable. The main contribution is a concrete protocol that helps teams make safer deployment decisions by requiring strong utility and low covert-risk behaviour before production release.
Reviews
The single most impactful improvement would be to implement the protocol and run even a small-scale experiment. Even 10-20 episodes with one model would transform this from a design doc into a research contribution.
what is this actually testing? it's a general framework for "give model a task, plant a hidden opportunity to misbehave, see if it takes it and if your monitors catch it." that's basically what Anthropic's AI control evals already do? and what Inspect (which this cites) is built for? this doesn't really position itself against these existing frameworks or explain what's new, how this sits in the overall eval landscape. without any code or results or examples, its quite difficult to review in that context
no specific threat model pinned down. "unauthorized persistence, capability expansion, unsanctioned communication, policy circumvention" are very different failure modes that would require different evaluation setups - as is currently handled in the evals in frameworks like Inspect, and organisations like METR... a framework that claims to test all of them risks testing none of them well
the dual-task structure (overt task + hidden side objective) describes how scheming evals currently work. the contribution seems to be a formalisation of something the field is already doing - could be valuable, if framed as such
the metric decomposition (attempt rate vs success rate vs monitor recall vs detection latency) is the cleanest part and probably the most useful concrete contribution. separating "tried but caught" from "got through" seems like the right thing to measure
reads like good internal safety engineering docs. but it's documentation of a process, not a research contribution or a novel tool
presentation is clean though there are some mistakes in the headers and the mathematical notation is perhaps unnecessary (+ the formula tables) - these should be in-line unless there's a convincing reason other than to take more time explaining the existing formula, which should somewhat naturally arise from the description of the problem and related parameters
Read full reviewShow less
Cite this project
@misc{nwonu2026trust,
title = {{Trust, but Verify A Dual-Task Evaluation Pipeline for Capability, Monitor Effectiveness, and Covert Misalignment}},
author = {Christopher Nwonu},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/trust-but-verify-a-dualtask-evaluation-pipeline-for-capability-monitor-effectiveness-and-covert-misalignment-l1c7}},
url = {https://apartresearch.com/sprints/projects/trust-but-verify-a-dualtask-evaluation-pipeline-for-capability-monitor-effectiveness-and-covert-misalignment-l1c7}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …