Axiom Zero: AlphaZero-Style Reinforcement Learning for Automated Formal Verification of Python Programs
Mufaro Rukuni, Brain Monzora · Team Axiom_Zer0
Submitted to The Secure Program Synthesis Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
The rapid growth of AI-generated code has created a verification crisis: Gross World Lines of Code (LoC) is expanding at unprecedented rates, yet developers have no systematic guarantee that the code produced by large language models behaves as intended. We present Axiom Zero, a compiler and reinforcement learning system that translates Python/PyTorch source code into formal proofs verified by the Lean 4 proof assistant, using an AlphaZero-style agent to discover those proofs automatically without human-labelled examples. The system comprises four implemented phases: (1) a parse pipeline that converts Python source into a normalized intermediate representation with type and tensor-shape analysis; (2) a proof environment modelling theorem proving as a two-player game with 39 curated tactics across 10 categories; (3) a policy-plus-value network with MCTS tree search trained via self-play; and (4) a Python-to-Lean 4 compiler with difficulty-stratified proof-hole filling. All 143 unit tests pass across 3,450 lines of dependency-free Python. The Lean 4 kernel serves as the binary oracle: a proof either compiles or it does not, providing a clean reward signal that eliminates the need for human annotation. Axiom Zero demonstrates that the game-theoretic self-play paradigm, previously applied to chess and Go, can be transferred to the domain of program verification, offering a scalable path toward trustworthy AI-assisted software development.

Reviews
The write-up for this project was clear, well-polished communication. I would be very interested to see the results of Phase 5, which are ongoing. The 3.5 design decisions all seemed appropriate and improved usability. The test suite had good, relevant and broad coverage.
I was not convinced by claim of 'isomorphism' or even substantial similarity to superhuman Go via self-play - the prover and kernel do not coevolve, and this approach is more aptly described as an expert iteration / single-agent search method.
I think limiting it to Lean 4 is basically fine.
This is a very ambitious project and it's great to see the authors got very far! A larger evaluation with more test cases would be useful next -- and, ideally, the authors would verify the correctness of their translation steps.
Cite this project
@misc{rukuni2026axiom,
title = {{Axiom Zero: AlphaZero-Style Reinforcement Learning for Automated Formal Verification of Python Programs}},
author = {Mufaro Rukuni and Brain Monzora},
year = {2026},
month = may,
note = {Submitted to The Secure Program Synthesis Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/axiom-zero-alphazerostyle-reinforcement-learning-for-automated-formal-verification-of-python-programs-aqr7}},
url = {https://apartresearch.com/sprints/projects/axiom-zero-alphazerostyle-reinforcement-learning-for-automated-formal-verification-of-python-programs-aqr7}
}More from The Secure Program Synthesis Hackathon
- View project: Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Lida Safety
Specifications for real systems do not exist as one-shot artifacts: the user's intent emerges as they discover edge cases, rewrite drafts, and react to failing tests. We present an iterative pipeline that takes this …
- View project: AgentSpecGap
AgentSpecGap
solo-team
This prototype extracts rules from system prompts, tool descriptions, and runtime config. Rules are classified into one of interface validation, authorization check, workflow ordering validation, runtime validation, …
- View project: SpecGap Arena
SpecGap Arena
Obligation Cartographers
SpecGap Arena is a benchmark and framework that exposes how incomplete specifications let plausible but incorrect code pass public tests. It synthesizes missing semantic obligations (security boundaries, invariants, …