Decoding AI Manipulation
We study and evaluate manipulation risks in frontier AI systems. We run applied research programs to bring global talent to AI safety.
Featured research
All researchNeurIPS 2025
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
ICLR 2025 · Oral
DarkBench: Benchmarking Dark Patterns in Large Language Models
Cooperative AI Foundation 2025 · Technical report
Multi-Agent Risks from Advanced AI
Pivot to AI Safety
Anyone can join a Sprint. Win or stand out, and you’re invited to Apart Studio, the way into the Apart Fellowship. Partnered Fellowships take direct applications.
How our programs connectWork with Apart
Independent evidence for consequential AI decisions
Frontier AI labs
Targeted model-behavior evaluations, independent testing of safeguards, and evidence for internal safety decisions.
Policymakers and public institutions
Empirical studies tied to policy questions, reproducible methods, and technical reporting.
AI safety researchers
Evaluation methods, benchmarks, risk frameworks, and collaborations that strengthen the field’s shared evidence base.
We also work with domain experts, builders and entrepreneurs, sponsors, and donors who want to make AI safer.
Contact us








