Skip to content
Sprint projectOct 27, 2024
1st place by peer review

Robust Machine Unlearning for Dangerous Capabilities

Neel Jay, Austin Meek, Joshua Ehi · Team Robust Machine Unlearning for Dangerous Capabilities

Submitted to AI Policy Hackathon at Johns Hopkins University. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Robust Machine Unlearning for Dangerous Capabilities

Presentation

Presentation: Robust Machine Unlearning for Dangerous Capabilities

Share

We test different unlearning methods to make models more robust against exploitation by malicious actors for the creation of bioweapons.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. On layout: follow template; On content: very focused to a clear problem, would have been most impfactful with the inclusion of clarity on types of nefarious users and their particular imputs to address variation of inputs and type of user

  2. Very unique and novel problem to approach. The UI/UX took points off due to no representation or example of usecase or users. Teamwork took points since commits came from one user.

  3. Novel problem and solid approach, and the technical details shows off a strong understanding of the relevant research question. However, the submission is more like a paper than a technical solution. Adding a video demonstration and showing the UI would be very helpful for understanding

  4. I would to get information for my own learners

Cite this project

@misc{jay2024robust,
  title = {{Robust Machine Unlearning for Dangerous Capabilities}},
  author = {Neel Jay and Austin Meek and Joshua Ehi},
  year = {2024},
  month = oct,
  note = {Submitted to AI Policy Hackathon at Johns Hopkins University, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/robust-machine-unlearning-for-dangerous-capabilities}},
  url = {https://apartresearch.com/sprints/projects/robust-machine-unlearning-for-dangerous-capabilities}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026