TheWizard
Guy Nachshon, Zack Zornstain · Team Oz Labs
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
TheWizard is a tiny cyber world model built to defend systems against AI agents. TheWizard constructs a symbolic digital twin of an environment and simulates how agents plan and act—files touched, logs queried, credentials accessed—to generate 10–20 plausible futures from any partial trajectory. If an agent’s predicted path contains harmful intent, TheWizard blocks it before execution. Powered by HRM-, TRM-, and VibeThinker-inspired reasoning enhancements, our 10–20M model brings anticipatory defense, agent-misuse detection, and multi-step threat prevention to edge devices.
Reviews
Interesting idea! Would be keen to see how this works in cases where the user genuinely wants to perform actions which are suspicious. The current setup seems to indicate that it might block those requests all the time, making it frustrating for the user, and more likely they would turn it off.
As a proof-of-concept, TheWizard shows promising groundwork.
It would help to include a brief note about operational cost. How often does Wizard need to run predictions, and how does that scale with agent complexity? Clarifying this would help assess compute cost vs. benefit.
I would also like to see a clearer articulation of the deployment model. Who runs the digital twin and where does it sit relative to the agent? A bit more framing around the types of high-stakes domains this is intended for would make the real-world value even more concrete.
Cite this project
@misc{nachshon2025thewizard,
title = {{TheWizard}},
author = {Guy Nachshon and Zack Zornstain},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/thewizard-u2dk}},
url = {https://apartresearch.com/sprints/projects/thewizard-u2dk}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …