Adding Document Summaries to Control Arena
Keshav Shenoy
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
One of the current focuses of AI Control research is finding new settings within which to examine the control problem. In this hackathon, we looked at defining a new AI control setting based around summarizing documents to mislead an overseer. We expect this problem to become more important as AI capabilities are more frequently used for synthesizibg and summarizing information, especially in AI safety research. As such, we worked on adding a new setting to the UKAISI Control Arena repository. The current state is a functional Inpect eval where redteams and blueteams could iterate on different strategies. However, the setting is not integrated into the Control Arena framework and only the most basic redteam & blueteam strategies have been implmeneted. Future work would move protocol and policy code into the existing Control Arena baseline, enable native pdf support, iterate further adversarially, and streamline AI calls which can be made asynchronously.
Reviews
No public critique yet.
Cite this project
@misc{shenoy2025adding,
title = {{Adding Document Summaries to Control Arena}},
author = {Keshav Shenoy},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/adding-document-summaries-to-control-arena-fggr}},
url = {https://apartresearch.com/sprints/projects/adding-document-summaries-to-control-arena-fggr}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …