
Feb 10 - 13, 2023Online and in person
Scale Oversight for Machine Learning Hackathon
This Sprint has ended.
- Sign-ups
- 24
Projects from this Sprint are not published on the site.
Join us for the fifth Alignment Jam where we get to spend 48 hours of intense research on how we can measure and monitor the safety of large-scale machine learning models. Work on safety benchmarks, models detecting faults in other models, self-monitoring systems , and so much else!
Overview
Hosted by Esben Kran, pseudobison, Zaki, fbarez, ruiqi-zhong · #alignmentjam
Join us for the fifth Alignment Jam where we get to spend 48 hours of intense research on how we can measure and monitor the safety of large-scale machine learning models. Work on safety benchmarks, models detecting faults in other models, self-monitoring systems , and so much else!
🏆$2,000 on the line

Measuring and monitoring safety
To make sure large machine learning models follow what we want them to do, we have to have people monitoring their safety. BUT, it is indeed very hard for just one person to monitor all the outputs of ChatGPT...
The objective of this hackathon is to research scalable solutions to this problem!
- Can we create good benchmarks that run independently of human oversight?
- Can we train AI models themselves to find faults in other models?
- Can we create ways for one human to monitor a much larger amount of data?
- Can we reduce the misgeneralization of the original model using some novel method?
These are all very interesting questions that we're excited to see your answers for during theses 48 hours
Reading group
Join the Discord above to be a part of the reading group where we read up on the research within scaling oversight! The current pieces are:
Resources
Inspiration
Inspiring resources for scalable oversight and ML safety:
- This lecture explains how future machine learning and AI systems might look and how we might predict emergent behaviour from large systems, something that is increasingly important in the context of scalable oversight: YouTube video link
- Watch Cambridge professor David Krueger's talk on ML safety: YouTube video link
- Watch the Center for AI Safety's Dan Hendrycks' short (10m) lecture on transparency in machine learning: YouTube video link
- Watch this lecture on Trojan neural networks, a way to study when neural networks diverge from our expectations: YouTube video link
Get notified when the intro talk stream starts on the Friday of the event!
Scale Oversight resources
Join us for the fifth Alignment Jam where we get to spend 48 hours of intense research on how we can measure and monitor the safety of large-scale machine learning models. Work on safety benchmarks, models detecting faults in other models, self-monitoring systems, and so much else!
To make sure large machine learning models follow what we want them to do, we have to have people monitoring their safety. BUT, it is indeed very hard for just one person to monitor all the outputs of ChatGPT...
The objective of this hackathon is to research scalable solutions to this problem!
- Can we create good benchmarks that run independently of human oversight?
- Can we train AI models themselves to find faults in other models?
- Can we create ways for one human to monitor a much larger amount of data?
- Can we reduce the misgeneralization of the original model using some novel method?
These are all very interesting questions that we're excited to see your answers for during theses 48 hours!
Dive deeper:
- Measuring Progress on Scalable Oversight for Large Language Models
- SafeBench competition example ideas
- Benchmarks in AI safety by Isabella Duan
Use this API key for OpenAI API access: [API key removed]
You probably want to view this website on a computer or laptop.
See here how to upload your project to the hackathon page and copy the PDF report template here.
See more resources here.
Local sites
Aarhus Scale Oversight Hackathon
effective-altruism-denmark
Join in the Nobelpark at Aarhus University for 48 hours of fun research!
Event page: Aarhus Scale Oversight Hackathon (opens in new tab)Copenhagen Scale Oversight Hackathon
effective-altruism-denmark
Join us from DTU, ITU and KU for the machine learning research hackathon with Alignment Jams! Howitzvej 30, 2000 Frederiksberg
Event page: Copenhagen Scale Oversight Hackathon (opens in new tab)EA Tech London x Safe AI London: Supercharging Alignment Hackathon
We'll be hacking away at the LISA office all weekend, come and join us!
Event page: EA Tech London x Safe AI London: Supercharging Alignment Hackathon (opens in new tab)Stanford x Berkeley AI Oversight Hackathon
Stanford AI Alignment and Berkeley AI Safety Student Initiative are hosting a joint AI alignment hackathon!
Event page: Stanford x Berkeley AI Oversight Hackathon (opens in new tab)Virtual Scale Oversight Hackathon
Join with teams online in the great virtual hackathon space!
Event page: Virtual Scale Oversight Hackathon (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com