Red Teaming A Narrow Path: ControlAI Policy Sprint by Aritra Das and Vaani Goenka
Vaani Goenka, Aritra Das · Team Controllers Of Ai
Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This research analyses two proposed AI governance policies – prohibiting recursive self-improvement in AI systems and mandating safety cases for deployment – through historical precedent analysis, agent-based modeling, and formal verification. Examining failures in analogous regulations (Basel II, BWC, NSG, Wassenaar, SEC rules), we identify systemic vulnerabilities: definitional ambiguity, enforcement leakage, and the impossibility of exhaustively enumerating safety scenarios for advanced AI. Agent-based simulations demonstrate that even minor probabilities of enforcement algorithm leakage or successful system ‘relabeling’ lead inevitably to policy failure, with evasion skill rapidly outpacing controls. Testing with current SoTA LLMs reveals inherent knowledge leakage risks, while formal analysis proves the fundamental contradiction in assuming finite safety cases for superintelligent systems. Our findings urge policymakers to prioritize precise definitions, acknowledge inherent verification limits for Ais improving Ais, and develop dynamic, leakage-resistant enforcement mechanisms, recognizing that proposed controls offer incomplete solutions against determined circumvention.
Reviews
Very interesting and novel approach, enjoyed reading it.
The concrete proposals seem broadly quite sensible.
Would like to see these researchers continue to apply themselves to policy.
Well-written, and correctly identifies many of the potential downsides of A Narrow Path. If restrictions on open source AI are necessary to prevent superintelligence for 20 years, we believe this is an acceptable cost, and indeed we clearly say that models that fall under the licensing regime must not be open sourced.
Seems reasonable that Narrow Path should point to a definition of what an environment is
Some of the risks identified, such as that of regulatory capture I agree are important - and while we took measures to try to address them, not completely satisfactory.
Overall, this is a thoughtful critique. The Basel II example is a realllly good one to use with policymakers and I am immediately thinking about places to use it. I need to learn more, similarly, about the SEC kill-switch example you raise -- it sounds fascinating but I'm previously unfamiliar with the case.
I liked the agent-based modeling as well, though think it could be have been helpful to provide clearer "so what" impacts of the analysis.
This is a great project. I'd be excited to see these researchers continue to apply themselves to problems in AI policy.
Cite this project
@misc{goenka2025red,
title = {{Red Teaming A Narrow Path: ControlAI Policy Sprint by Aritra Das and Vaani Goenka}},
author = {Vaani Goenka and Aritra Das},
year = {2025},
month = jun,
note = {Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/red-teaming-a-narrow-path-controlai-policy-sprint-by-aritra-das-and-vaani-goenka-w5w9}},
url = {https://apartresearch.com/sprints/projects/red-teaming-a-narrow-path-controlai-policy-sprint-by-aritra-das-and-vaani-goenka-w5w9}
}More from Red Teaming A Narrow Path: ControlAI Policy Sprint
- View project: Treaty Enforcement in China
Treaty Enforcement in China
JackAI
This report red-teams A Narrow Path’s international treaty proposal by stress-testing its assumptions in the Chinese context. It identifies key failure modes—regulatory capture, compute-based loopholes, and covert …
- View project: Four Paths to Failure: Red Teaming ASI Governance
Four Paths to Failure: Red Teaming ASI Governance
Shoggoth Prevention Squad
We stress‑tested A Narrow Path Phase 0—the proposed 20‑year moratorium on training artificial super‑intelligence (ASI)—during a one‑day red‑teaming hackathon. Drawing on rapid literature reviews, historical analogues …
- View project: Moratorium on the development of general AI systems
Moratorium on the development of general AI systems
G_control
All six policies are red teamed step-by-step systematically. We initially corrected vague definitions and also found that the policies regarding the capabilities of AI systems lack technical soundness and that more …