SafetyBench
Hendrik Wirthwein · Team Berlin-hacker
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Recent work demonstrating strategic deception in LLM agents playing Diplomacy and scenario forecasting efforts like AI 2027 highlight the urgent need for controlled frameworks to evaluate multi-agent AI behaviors systematically. We introduce **apart**, a hybrid multi-agent orchestration framework combining configuration-driven game mechanics with LLM-powered simulation engines, enabling systematic comparison of heterogeneous agent-model combinations (Google Gemini, OpenAI, local Ollama models) across configurable scenarios while maintaining computational efficiency and reproducibility. The framework incorporates post-hoc behavioral analysis detecting dangerous patterns (power-seeking, deception, rule exploitation) through lightweight signal collection and LLM-as-judge evaluation, providing quantitative safety scoring with minimal performance impact (<5ms per step).
Reviews
No public critique yet.
Cite this project
@misc{wirthwein2025safetybench,
title = {{SafetyBench}},
author = {Hendrik Wirthwein},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/safetybench-xh8g}},
url = {https://apartresearch.com/sprints/projects/safetybench-xh8g}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …