An Exploration of Current Theory of Mind Evals
John Henderson, Alan Fung, Bachar Moustapha · Team AgentToM
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
We evaluated the performance of a prominent large language model from Anthropic, on the Theory of Mind evaluation developed by the AI Safety Institute (AISI). Our investigation revealed issues with the dataset used by AISI for this evaluation.
Reviews
Great that you identified the false negatives and performed further manual assessments on these (and that you opened the issue on the AISI repo!). Including some background on the nature of the dataset used might have improved readability. While the work is valuable, I would also have liked to see more detail on future research directions.
I’m glad that someone did this work, and it’s useful to check and point out errors in important datasets used by AISI. I think however that we could learn more from exploring new deception-detection techniques on small examples than from just running a particular model on a particular dataset.
Cite this project
@misc{henderson2024exploration,
title = {{An Exploration of Current Theory of Mind Evals}},
author = {John Henderson and Alan Fung and Bachar Moustapha},
year = {2024},
month = jul,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/an-exploration-of-current-theory-of-mind-evals}},
url = {https://apartresearch.com/sprints/projects/an-exploration-of-current-theory-of-mind-evals}
}More from Deception Detection Hackathon: Preventing AI deception
- 1st place by peer reviewView project: Sandbag Detection through Model Degradation
Sandbag Detection through Model Degradation
Team Truth Serum
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
- View project: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DeceptionRepE
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …