Robust Machine Unlearning for Dangerous Capabilities
Neel Jay, Austin Meek, Joshua Ehi · Team Robust Machine Unlearning for Dangerous Capabilities
Submitted to AI Policy Hackathon at Johns Hopkins University. Sprint projects are early-stage work by participants, not Apart Research publications.
We test different unlearning methods to make models more robust against exploitation by malicious actors for the creation of bioweapons.
Reviews
On layout: follow template; On content: very focused to a clear problem, would have been most impfactful with the inclusion of clarity on types of nefarious users and their particular imputs to address variation of inputs and type of user
Very unique and novel problem to approach. The UI/UX took points off due to no representation or example of usecase or users. Teamwork took points since commits came from one user.
Novel problem and solid approach, and the technical details shows off a strong understanding of the relevant research question. However, the submission is more like a paper than a technical solution. Adding a video demonstration and showing the UI would be very helpful for understanding
I would to get information for my own learners
Cite this project
@misc{jay2024robust,
title = {{Robust Machine Unlearning for Dangerous Capabilities}},
author = {Neel Jay and Austin Meek and Joshua Ehi},
year = {2024},
month = oct,
note = {Submitted to AI Policy Hackathon at Johns Hopkins University, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/robust-machine-unlearning-for-dangerous-capabilities}},
url = {https://apartresearch.com/sprints/projects/robust-machine-unlearning-for-dangerous-capabilities}
}More from AI Policy Hackathon at Johns Hopkins University
- View project: Infectious Disease Outbreak Prediction and Dashboard
Infectious Disease Outbreak Prediction and Dashboard
Infectious Diseace Dashboards
Our project developed an interactive dashboard to monitor, visualize, and analyze infectious disease outbreaks worldwide. It consolidates historical data from sources like WHO, OWID, and CDC for diseases including …
- View project: Modernizing DC’s Emergency Communications
Modernizing DC’s Emergency Communications
AI-CAD
The District of Columbia proposes implementing an AI-enabled Computer-Aided Dispatch (CAD) system to address critical deficiencies in our current emergency alert infrastructure. This policy establishes a framework for …
- View project: Improving Llama-3-8b Hallucination Robustness in Medical Q&A Using Feature Steering
Improving Llama-3-8b Hallucination Robustness in Medical Q&A Using Feature Steering
Gradients Anatomy
This paper addresses hallucinations in large language models (LLMs) within critical domains like medicine. It proposes and demonstrates methods to: Reduce Hallucination Probability: By using Llama-3-8B-Instruct and its …