Kernel of Trust: Evaluating AI Control Protocols Using Open-Source Data
AJ Weeks · Team Stockholm Center for AI Safety (SCAIS)
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
We propose a new setting to evaluate AI Control policies, namely the change history of the Linux kernel. This real-world data source provides a complement to the APPS backdooring setting used in the original AI Control paper.
Reviews
No public critique yet.
Cite this project
@misc{weeks2025kernel,
title = {{Kernel of Trust: Evaluating AI Control Protocols Using Open-Source Data}},
author = {AJ Weeks},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/kernel-of-trust-evaluating-ai-control-protocols-using-opensource-data-148y}},
url = {https://apartresearch.com/sprints/projects/kernel-of-trust-evaluating-ai-control-protocols-using-opensource-data-148y}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …