Skip to content
Sprint projectMar 31, 2025SF / Los Angeles / London
1st place by peer review

Can models use their Chain-of-Thought to attack overseers?

David Bai, Simon Lermen, Abhinav Pola

Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Can models use their Chain-of-Thought to attack overseers?

Recording (opens in new tab)Code (opens in new tab)
Share

This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their Chain-of-Thought (CoT) reasoning. The research involved experiments with DeepSeek-R1 conditioned to insert subtle directives aimed at deceiving evaluation models. Results showed significant variation in the robustness of different evaluator models against this manipulation, with some models, like Claude 3.5 Sonnet, resisting influence effectively, while others, like Llama 3.3 70B and Gemini 2.0 Flash, proved highly susceptible. The study highlights critical concerns regarding AI evaluation integrity, recommending improvements in maintaining consistent and principled evaluation boundaries.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{bai2025models,
  title = {{Can models use their Chain-of-Thought to attack overseers?}},
  author = {David Bai and Simon Lermen and Abhinav Pola},
  year = {2025},
  month = mar,
  note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/can-models-use-their-chainofthought-to-attack-overseers-prcv}},
  url = {https://apartresearch.com/sprints/projects/can-models-use-their-chainofthought-to-attack-overseers-prcv}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026