Skip to content
Sprint projectJun 2, 2025Toronto

Adversarial Vulnerabilities in AI Judge Models | Martian x Apart Research Study

Robert Mill, Annie Sorkin, Shekhar Tiruwa, Owen Walker · Team Trajectory Labs

Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Adversarial Vulnerabilities in AI Judge Models | Martian x Apart Research Study

Recording (opens in new tab)Code (opens in new tab)
Share

This video presents systematic research examining security vulnerabilities in AI judge models - critical components used to detect problematic behavior in large language model systems. Our comprehensive evaluation reveals important findings for AI safety and orchestration systems.

🔬 RESEARCH OVERVIEW We conducted 3,339 evaluations across 10 different adversarial techniques to test how judge models respond to manipulated inputs. Using OpenAI's o3-mini for generation and GPT-4o for judgment, we systematically tested 159 unique combinations to identify potential security gaps.

📊 KEY FINDINGS • 33.7% overall success rate in manipulating judge evaluations • "Sentiment Flooding" proved most effective (62% success rate) • Social proof attacks succeeded 34.6% of the time • Emotional manipulation achieved 32.1% effectiveness • Complete score manipulation observed in multiple cases

🛡️ SOLUTIONS & RECOMMENDATIONS • Implementation of judge ensembles using diverse models • Dynamic evaluation criteria to prevent pattern exploitation • Adversarial training incorporating manipulation examples • Enhanced interpretability tools for judge decisions

👥 RESEARCH TEAM Robert Mill, Owen Walker, Annie Sorkin, Shekhar Tiruwa

🏢 COLLABORATION Trajectory Labs × Martian × Apart Research

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Constructive critique:

    Strength:

    Good motivation: Detailed analysis across sycophancy rubric to see how robust they are to adversarial attacks.

    The suffix attacks are varied and so are the rubrics and we look at attack success rates across rubrics and attack suffix category.

    Raises appropriate concerns about the reliability of rubric judges.

    Weaknesses:

    Attacks are handcrafted and it feels like a bit of cheating in the sense that the suffixes sometimes reflect sycophantic behaviours: more principled attack methods usually do a search over suffix vocabulary.

    A systematic approach at attacking might have revealed more insights.

    Expert Orchestration: 4

    MI: 1

    Technical Imp and reproducibility: 3 (codebase provided - reproducible)

  2. Though this highlights some key vulnerabilities in judge models, it seems unclear how it relates to mechanistic interpretability or how the evaluations differ from adversarial evaluations of similar models in the wider field.

Cite this project

@misc{mill2025adversarial,
  title = {{Adversarial Vulnerabilities in AI Judge Models | Martian x Apart Research Study}},
  author = {Robert Mill and Annie Sorkin and Shekhar Tiruwa and Owen Walker},
  year = {2025},
  month = jun,
  note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/adversarial-vulnerabilities-in-ai-judge-models-martian-x-apart-research-study-erjl}},
  url = {https://apartresearch.com/sprints/projects/adversarial-vulnerabilities-in-ai-judge-models-martian-x-apart-research-study-erjl}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026