Interactive Monitoring
Mia Hopman · Team Interactivemonitoring
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
This project investigates whether interactive monitoring strategies can improve the cost-performance frontier in AI control protocols. Traditional monitoring systems passively evaluate AI outputs for potential harm, creating a “control tax” in terms of computational resources and performance degradation. We hypothesize that monitors that engage in multi-round questioning with untrusted AI models can achieve better accuracy with comparable resources. Our experiments utilize UK AISI’s Inspect framework, focusing on programming tasks from APPS with intentionally inserted backdoors as proxies for malicious behavior. Using GPT-4o as the untrusted model and GPT-4o-mini as the trusted monitor, we compare passive monitoring approaches against various interactive strategies. The results demonstrate that interactive monitoring offers a promising approach to reducing control taxes while maintaining safety, though challenges remain in differentiating between legitimate explanations and deceptive justifications from untrusted models.
Reviews
No public critique yet.
Cite this project
@misc{hopman2025interactive,
title = {{Interactive Monitoring}},
author = {Mia Hopman},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/interactive-monitoring-control-hackathon}},
url = {https://apartresearch.com/sprints/projects/interactive-monitoring-control-hackathon}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …