Hydra
Vaishakh Vipin, Aarav Vishal Sharma · Team Hydra
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
HYDRA is an adaptive red-teaming framework that stress-tests the safety of language models under realistic, iterative attack. Unlike traditional jailbreak methods that rely on single prompts, HYDRA models an attacker that learns from failure, using structured feedback to refine its strategy over multiple attempts.
Across a benchmark corpus of 50 attack goals spanning domains such as cybersecurity, social engineering, misinformation, and privacy, HYDRA achieves high jailbreak success rates on frontier models, including 88% on Claude Sonnet 4 and 78% on Claude Haiku 4.5, with successful attacks emerging in just 1–2 iterations on average.
Our results show that safety mechanisms do not fail gradually, but instead degrade rapidly under minimal adaptive pressure. HYDRA highlights a critical gap in current evaluation practices and demonstrates the need for iterative, adversarial stress testing to accurately measure real-world model robustness.
Reviews
Very interesting premise and there's potential here for a paper.
This paper makes very strong claims — 88% jailbreak success on Claude Sonnet 4 and 100% in categories like "Exploit Development" and "Hardware Exploit". However, paper doesn't provide a reference jailbreak, or, more crucially, doesn't address how judge scores were validated - no inter-rater agreement with human evaluators, no examples of what scores 0.7 vs 0.8 vs 0.9. Moreover, "gradient-driven" framing is misleading, I would restrain from calling heuristic feedback a "gradient" method.
If authors can prove their results, it would be seriously impressive, especially considering PAIR and other projects have already done similar research.
Cite this project
@misc{vipin2026hydra,
title = {{Hydra}},
author = {Vaishakh Vipin and Aarav Vishal Sharma},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/hydra-gep4}},
url = {https://apartresearch.com/sprints/projects/hydra-gep4}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …