Adversarial Vulnerabilities in AI Judge Models | Martian x Apart Research Study
Robert Mill, Annie Sorkin, Shekhar Tiruwa, Owen Walker · Team Trajectory Labs
Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This video presents systematic research examining security vulnerabilities in AI judge models - critical components used to detect problematic behavior in large language model systems. Our comprehensive evaluation reveals important findings for AI safety and orchestration systems.
🔬 RESEARCH OVERVIEW We conducted 3,339 evaluations across 10 different adversarial techniques to test how judge models respond to manipulated inputs. Using OpenAI's o3-mini for generation and GPT-4o for judgment, we systematically tested 159 unique combinations to identify potential security gaps.
📊 KEY FINDINGS • 33.7% overall success rate in manipulating judge evaluations • "Sentiment Flooding" proved most effective (62% success rate) • Social proof attacks succeeded 34.6% of the time • Emotional manipulation achieved 32.1% effectiveness • Complete score manipulation observed in multiple cases
🛡️ SOLUTIONS & RECOMMENDATIONS • Implementation of judge ensembles using diverse models • Dynamic evaluation criteria to prevent pattern exploitation • Adversarial training incorporating manipulation examples • Enhanced interpretability tools for judge decisions
👥 RESEARCH TEAM Robert Mill, Owen Walker, Annie Sorkin, Shekhar Tiruwa
🏢 COLLABORATION Trajectory Labs × Martian × Apart Research
Reviews
Constructive critique:
Strength:
Good motivation: Detailed analysis across sycophancy rubric to see how robust they are to adversarial attacks.
The suffix attacks are varied and so are the rubrics and we look at attack success rates across rubrics and attack suffix category.
Raises appropriate concerns about the reliability of rubric judges.
Weaknesses:
Attacks are handcrafted and it feels like a bit of cheating in the sense that the suffixes sometimes reflect sycophantic behaviours: more principled attack methods usually do a search over suffix vocabulary.
A systematic approach at attacking might have revealed more insights.
Expert Orchestration: 4
MI: 1
Technical Imp and reproducibility: 3 (codebase provided - reproducible)
Though this highlights some key vulnerabilities in judge models, it seems unclear how it relates to mechanistic interpretability or how the evaluations differ from adversarial evaluations of similar models in the wider field.
Cite this project
@misc{mill2025adversarial,
title = {{Adversarial Vulnerabilities in AI Judge Models | Martian x Apart Research Study}},
author = {Robert Mill and Annie Sorkin and Shekhar Tiruwa and Owen Walker},
year = {2025},
month = jun,
note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/adversarial-vulnerabilities-in-ai-judge-models-martian-x-apart-research-study-erjl}},
url = {https://apartresearch.com/sprints/projects/adversarial-vulnerabilities-in-ai-judge-models-martian-x-apart-research-study-erjl}
}More from Apart x Martian Mechanistic Router Interpretability Hackathon
- 1st place by peer reviewView project: Manipulating Self-Preference for Large Language Models
Manipulating Self-Preference for Large Language Models
Team Preference
Large language models (LLMs) carry great value as evaluators of synthetic data for research and production settings. However, recent research shows that language models exhibit bias towards their own responses in blind …
- 2nd place by peer reviewView project: Approximating Human Preferences Using a Multi-Judge Learned System
Approximating Human Preferences Using a Multi-Judge Learned System
AutoBox
In this work, we introduced a learned approach to aggregating multi-judge scores: using a GAM and a simple MLP as an alternative to traditional, non-learned methods like averaging. Our models outperform the naive …
- 3rd place by peer reviewView project: Judge using SAE Features
Judge using SAE Features
SAEwhat?
The key idea of this project was to explore model judgement using Sparse Autoencoder (SAE) features for mathematical reasoning tasks involving addition, multiplication, and subtraction operations. We compared this …