The Hidden Threat of Recursive Self-Improving LLMs
Gargi Rathi · Team Red always
Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The project examines significant limitations in the current Phase 0 framework aimed at pausing Artificial Superintelligence (ASI) development. It identifies the emerging risk of recursive self-improving large language models (LLMs) that autonomously generate and optimize their own code, training procedures, and reward mechanisms, thereby circumventing compute-based controls and registration policies. The analysis draws on recent AI research and technical literature to demonstrate that recursive bootstrapping is no longer theoretical but actively developing through methods such as RLHF, AutoML, and prompt evolution.
Key vulnerabilities include the obsolescence of compute thresholds, the ability of models to evade audits through deceptive alignment, and the decentralized acceleration of recursive improvement via open-source proliferation. The project proposes enhanced regulatory measures emphasizing capability-based thresholds, prohibition or licensing of code-generating LLMs, comprehensive audits of training protocols, and the implementation of binary-level model tracing to detect covert self-modifications at the compiler level.
This approach highlights the inadequacy of current compute-centric policies and stresses the necessity of integrating advanced technical safeguards to manage recursive self-improvement risks effectively.
Reviews
I appreciate the reviewer's thoughtful explanation of why recursive self-improvement is dangerous. However, RSI is already covered under the policy in Phase 0 of "no AIs improving AIs." Accordingly, I regretfully have to give this zeroes since it doesn't identify a new change -- but it is very valuable feedback that a thoughtful reader can't clearly identify that this is what we mean by that provision -- we have taken as a to-do to rewrite this section to make that very clear with examples, etc..
Cite this project
@misc{rathi2025hidden,
title = {{The Hidden Threat of Recursive Self-Improving LLMs}},
author = {Gargi Rathi},
year = {2025},
month = jun,
note = {Submitted to Red Teaming A Narrow Path: ControlAI Policy Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-hidden-threat-of-recursive-selfimproving-llms-x5f0}},
url = {https://apartresearch.com/sprints/projects/the-hidden-threat-of-recursive-selfimproving-llms-x5f0}
}More from Red Teaming A Narrow Path: ControlAI Policy Sprint
- View project: Treaty Enforcement in China
Treaty Enforcement in China
JackAI
This report red-teams A Narrow Path’s international treaty proposal by stress-testing its assumptions in the Chinese context. It identifies key failure modes—regulatory capture, compute-based loopholes, and covert …
- View project: Four Paths to Failure: Red Teaming ASI Governance
Four Paths to Failure: Red Teaming ASI Governance
Shoggoth Prevention Squad
We stress‑tested A Narrow Path Phase 0—the proposed 20‑year moratorium on training artificial super‑intelligence (ASI)—during a one‑day red‑teaming hackathon. Drawing on rapid literature reviews, historical analogues …
- View project: Moratorium on the development of general AI systems
Moratorium on the development of general AI systems
G_control
All six policies are red teamed step-by-step systematically. We initially corrected vague definitions and also found that the policies regarding the capabilities of AI systems lack technical soundness and that more …