Skip to content
Sprint projectJul 1, 2024

An Exploration of Current Theory of Mind Evals

John Henderson, Alan Fung, Bachar Moustapha · Team AgentToM

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: An Exploration of Current Theory of Mind Evals

Code (opens in new tab)
Share

We evaluated the performance of a prominent large language model from Anthropic, on the Theory of Mind evaluation developed by the AI Safety Institute (AISI). Our investigation revealed issues with the dataset used by AISI for this evaluation.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Great that you identified the false negatives and performed further manual assessments on these (and that you opened the issue on the AISI repo!). Including some background on the nature of the dataset used might have improved readability. While the work is valuable, I would also have liked to see more detail on future research directions.

  2. I’m glad that someone did this work, and it’s useful to check and point out errors in important datasets used by AISI. I think however that we could learn more from exploring new deception-detection techniques on small examples than from just running a particular model on a particular dataset.

Cite this project

@misc{henderson2024exploration,
  title = {{An Exploration of Current Theory of Mind Evals}},
  author = {John Henderson and Alan Fung and Bachar Moustapha},
  year = {2024},
  month = jul,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/an-exploration-of-current-theory-of-mind-evals}},
  url = {https://apartresearch.com/sprints/projects/an-exploration-of-current-theory-of-mind-evals}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026