Skip to content
Sprint projectJun 13, 2025Sonepat, Haryana, India

Red Teaming A Narrow Path: ControlAI Policy Sprint by Aritra Das and Vaani Goenka

Vaani Goenka, Aritra Das · Team Controllers Of Ai

Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Red Teaming A Narrow Path: ControlAI Policy Sprint by Aritra Das and Vaani Goenka

Share

This research analyses two proposed AI governance policies – prohibiting recursive self-improvement in AI systems and mandating safety cases for deployment – through historical precedent analysis, agent-based modeling, and formal verification. Examining failures in analogous regulations (Basel II, BWC, NSG, Wassenaar, SEC rules), we identify systemic vulnerabilities: definitional ambiguity, enforcement leakage, and the impossibility of exhaustively enumerating safety scenarios for advanced AI. Agent-based simulations demonstrate that even minor probabilities of enforcement algorithm leakage or successful system ‘relabeling’ lead inevitably to policy failure, with evasion skill rapidly outpacing controls. Testing with current SoTA LLMs reveals inherent knowledge leakage risks, while formal analysis proves the fundamental contradiction in assuming finite safety cases for superintelligent systems. Our findings urge policymakers to prioritize precise definitions, acknowledge inherent verification limits for Ais improving Ais, and develop dynamic, leakage-resistant enforcement mechanisms, recognizing that proposed controls offer incomplete solutions against determined circumvention.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the analysis realistically assess what government agencies, resources, and expertise would be needed to implement these policies? Are the identified implementation challenges specific and grounded in understanding of how similar policies have worked (or failed) in practice? Does the submission adequately consider bureaucratic, technical, and coordination complexities involved in enforcement? How well does the analysis account for real-world constraints like budget limitations, regulatory capture, and inter-agency coordination?

Does the analysis identify specific ways the policies could fail to prevent ASI development or be circumvented by determined actors? How thoroughly does the submission examine edge cases, loopholes, or unintended consequences that could undermine the 20-year goal? Does the assessment consider different threat models (state actors, rogue researchers, corporate actors) and how policies address each? Are the identified failure modes realistic and significant, or primarily theoretical edge cases?

Does the submission cite relevant historical examples of similar policies (nuclear non-proliferation, export controls, dual-use technology regulation) to support its arguments? Are claims backed by empirical data, documented case studies, or credible expert analysis rather than speculation? How well does the analysis draw lessons from comparable regulatory domains to assess likely outcomes? Does the submission avoid making unsupported assertions about what "would" or "could" happen without evidence?

  1. Very interesting and novel approach, enjoyed reading it.

    The concrete proposals seem broadly quite sensible.

    Would like to see these researchers continue to apply themselves to policy.

  2. Well-written, and correctly identifies many of the potential downsides of A Narrow Path. If restrictions on open source AI are necessary to prevent superintelligence for 20 years, we believe this is an acceptable cost, and indeed we clearly say that models that fall under the licensing regime must not be open sourced.

    Seems reasonable that Narrow Path should point to a definition of what an environment is

    Some of the risks identified, such as that of regulatory capture I agree are important - and while we took measures to try to address them, not completely satisfactory.

  3. Overall, this is a thoughtful critique. The Basel II example is a realllly good one to use with policymakers and I am immediately thinking about places to use it. I need to learn more, similarly, about the SEC kill-switch example you raise -- it sounds fascinating but I'm previously unfamiliar with the case.

    I liked the agent-based modeling as well, though think it could be have been helpful to provide clearer "so what" impacts of the analysis.

  4. This is a great project. I'd be excited to see these researchers continue to apply themselves to problems in AI policy.

Cite this project

@misc{goenka2025red,
  title = {{Red Teaming A Narrow Path: ControlAI Policy Sprint by Aritra Das and Vaani Goenka}},
  author = {Vaani Goenka and Aritra Das},
  year = {2025},
  month = jun,
  note = {Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/red-teaming-a-narrow-path-controlai-policy-sprint-by-aritra-das-and-vaani-goenka-w5w9}},
  url = {https://apartresearch.com/sprints/projects/red-teaming-a-narrow-path-controlai-policy-sprint-by-aritra-das-and-vaani-goenka-w5w9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026