Skip to content
Sprint projectJun 14, 2025Edinburgh UK

AI Assistance in AI alignment Improvement: Allow It!

Anthony Bailey · Team anthonybailey.net

Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: AI Assistance in AI alignment Improvement: Allow It!

Share

A Narrow Path currently includes a condition (A): “No AIs improving AIs” that underlies various parts of the document. It makes no exception for what I will abbreviate as AI^4: AI Assisting In AI Alignment Improvement. It should, because sufficiently many in AI safety while acknowledging its unique hazards still see value in exploring “AI helping with our alignment homework.” Specifically I am Red Teaming Phase 0 policy 3 (“Prohibit the development and use of AIs that improve other AIs”) arguing no policy that bans AI^4 will garner sufficient support to be agreed or enforced.

A variety of Deep Research implementations suggested 70-90% of those expressing relevant opinions in forums associated with AIXR concern would oppose such a ban.

If AI assistance in AI alignment improvements is to be allowed, that needs to be made more clear, and consequences to licensing and enforcement considered.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the analysis realistically assess what government agencies, resources, and expertise would be needed to implement these policies? Are the identified implementation challenges specific and grounded in understanding of how similar policies have worked (or failed) in practice? Does the submission adequately consider bureaucratic, technical, and coordination complexities involved in enforcement? How well does the analysis account for real-world constraints like budget limitations, regulatory capture, and inter-agency coordination?

Does the analysis identify specific ways the policies could fail to prevent ASI development or be circumvented by determined actors? How thoroughly does the submission examine edge cases, loopholes, or unintended consequences that could undermine the 20-year goal? Does the assessment consider different threat models (state actors, rogue researchers, corporate actors) and how policies address each? Are the identified failure modes realistic and significant, or primarily theoretical edge cases?

Does the submission cite relevant historical examples of similar policies (nuclear non-proliferation, export controls, dual-use technology regulation) to support its arguments? Are claims backed by empirical data, documented case studies, or credible expert analysis rather than speculation? How well does the analysis draw lessons from comparable regulatory domains to assess likely outcomes? Does the submission avoid making unsupported assertions about what "would" or "could" happen without evidence?

  1. Seems like a reasonable objection that is worth pondering and either changing or providing some justification for keeping as is.

    However, if changes to permit this allow for the possibility of an intelligence recursion, I don't think it should be allowed.

    Didn't really engage with the goal of preventing the development of superintelligence for 20+ years.

  2. We appreciate this thoughtful feedback. The argument about using AI cautiously for tasks to reduce AI risks is an important one to consider, and the arguments about the need to consider it for both principled and pragmatic reasons were interesting. It would have been helpful to see additional arguments about how likely it would be that a policy regime could succeed at correctly drawing the line between AI work that reduces AI risk and AI work that increases it, but nonetheless worth considering given, e.g., UK AISI's clear interest in this (as well as, of course, all major frontier AI companies)

Cite this project

@misc{bailey2025ai,
  title = {{AI Assistance in AI alignment Improvement: Allow It!}},
  author = {Anthony Bailey},
  year = {2025},
  month = jun,
  note = {Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ai-assistance-in-ai-alignment-improvement-allow-it-oru2}},
  url = {https://apartresearch.com/sprints/projects/ai-assistance-in-ai-alignment-improvement-allow-it-oru2}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026