Voyager - Self-Evolving AI Control Platform
Nav Sangameswaran, Preetham Sathyamurthy, Sabarish Varadharajan · Team Astroware Research
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Voyager is an AI Control Platform that treats AI safety as an engineering problem, not a policy afterthought. Today, frontier labs already use AI to accelerate AI capability, while safety work is still done by small human teams running occasional red-teaming exercises and reading papers. This is a losing game. Offensive use of AI is being automated and scaled, yet defensive use is fragmented and slow.
Voyager is our attempt to flip that. It turns cutting edge safety ideas such as evolutionary red-teaming, constitutional judging and mechanistic interpretability into a continuous control loop around real models that real organisations run. In this hackathon, we implemented the Offend stage of our OUDA lifecycle and exposed it as a SaaS product. A model owner connects their model, presses “audit”, and Voyager uses powerful models as adversaries, mutating psychologically rich attack strategies until they elicit failures such as deception, policy circumvention or harmful assistance. The system then writes an audit report that a risk or compliance team can act on.
In experiments, Voyager autonomously discovered non trivial deceptive behaviour in Claude Sonnet 4.5, one of the most safety tuned models available, with no hand written jailbreaks. That is the core def/acc story in one sentence: AI is already capable of doing serious safety work on AI, and wrapping that capability in a reusable control platform lets defence scale at something like the same rate as capability, for both frontier labs and enterprises.
Reviews
No public critique yet.
Cite this project
@misc{sangameswaran2025voyager,
title = {{Voyager - Self-Evolving AI Control Platform}},
author = {Nav Sangameswaran and Preetham Sathyamurthy and Sabarish Varadharajan},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/voyager-selfevolving-ai-control-platform-72m9}},
url = {https://apartresearch.com/sprints/projects/voyager-selfevolving-ai-control-platform-72m9}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …