Mar 10, 2025
Inspiring People to Go into RL Interp
Christopher Kinoshita, Siya Singh Deneille Guiseppi
Summary
This project is attempting to complete the Public Education Track, taking inspiration from ideas 1 and 4. The journey mapping was inspired by bluedot impact and aims to create a course that helps explain the need for work to be done in Reinforcement Learning (RL) interp, especially in the problems of reward hacking and goal misgeneralization. The point of the game is to make a humorous example of what could happen due to a lack of AI safety (not specifically goal misalignment or reward hacking) and is meant to be a fun introduction for nontechnical people to even care about AI safety.
Cite this work:
@misc {
title={
Inspiring People to Go into RL Interp
},
author={
Christopher Kinoshita, Siya Singh Deneille Guiseppi
},
date={
3/10/25
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}