AI Assistance in AI alignment Improvement: Allow It!
Anthony Bailey · Team anthonybailey.net
Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
A Narrow Path currently includes a condition (A): “No AIs improving AIs” that underlies various parts of the document. It makes no exception for what I will abbreviate as AI^4: AI Assisting In AI Alignment Improvement. It should, because sufficiently many in AI safety while acknowledging its unique hazards still see value in exploring “AI helping with our alignment homework.” Specifically I am Red Teaming Phase 0 policy 3 (“Prohibit the development and use of AIs that improve other AIs”) arguing no policy that bans AI^4 will garner sufficient support to be agreed or enforced.
A variety of Deep Research implementations suggested 70-90% of those expressing relevant opinions in forums associated with AIXR concern would oppose such a ban.
If AI assistance in AI alignment improvements is to be allowed, that needs to be made more clear, and consequences to licensing and enforcement considered.
Reviews
Seems like a reasonable objection that is worth pondering and either changing or providing some justification for keeping as is.
However, if changes to permit this allow for the possibility of an intelligence recursion, I don't think it should be allowed.
Didn't really engage with the goal of preventing the development of superintelligence for 20+ years.
We appreciate this thoughtful feedback. The argument about using AI cautiously for tasks to reduce AI risks is an important one to consider, and the arguments about the need to consider it for both principled and pragmatic reasons were interesting. It would have been helpful to see additional arguments about how likely it would be that a policy regime could succeed at correctly drawing the line between AI work that reduces AI risk and AI work that increases it, but nonetheless worth considering given, e.g., UK AISI's clear interest in this (as well as, of course, all major frontier AI companies)
Cite this project
@misc{bailey2025ai,
title = {{AI Assistance in AI alignment Improvement: Allow It!}},
author = {Anthony Bailey},
year = {2025},
month = jun,
note = {Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/ai-assistance-in-ai-alignment-improvement-allow-it-oru2}},
url = {https://apartresearch.com/sprints/projects/ai-assistance-in-ai-alignment-improvement-allow-it-oru2}
}More from Red Teaming A Narrow Path: ControlAI Policy Sprint
- View project: Treaty Enforcement in China
Treaty Enforcement in China
JackAI
This report red-teams A Narrow Path’s international treaty proposal by stress-testing its assumptions in the Chinese context. It identifies key failure modes—regulatory capture, compute-based loopholes, and covert …
- View project: Four Paths to Failure: Red Teaming ASI Governance
Four Paths to Failure: Red Teaming ASI Governance
Shoggoth Prevention Squad
We stress‑tested A Narrow Path Phase 0—the proposed 20‑year moratorium on training artificial super‑intelligence (ASI)—during a one‑day red‑teaming hackathon. Drawing on rapid literature reviews, historical analogues …
- View project: Moratorium on the development of general AI systems
Moratorium on the development of general AI systems
G_control
All six policies are red teamed step-by-step systematically. We initially corrected vague definitions and also found that the policies regarding the capabilities of AI systems lack technical soundness and that more …