Skip to content
Sprint projectMar 31, 2025London
3rd place by peer review

Model Models: Simulating a Trusted Monitor

Louis Thomson, Alex Holness-Tofts, Jasmijn Remmers, Xanthe Spence, Lennie Wells, Charlie Griffin · Team Charlie's Angels

Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Model Models: Simulating a Trusted Monitor

Code (opens in new tab)
Share

We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular codebase which we hope can be used to streamline future research into this question.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{thomson2025model,
  title = {{Model Models: Simulating a Trusted Monitor}},
  author = {Louis Thomson and Alex Holness-Tofts and Jasmijn Remmers and Xanthe Spence and Lennie Wells and Charlie Griffin},
  year = {2025},
  month = mar,
  note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/model-models-simulating-a-trusted-monitor-r682}},
  url = {https://apartresearch.com/sprints/projects/model-models-simulating-a-trusted-monitor-r682}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026