Skip to content
Sprint projectMay 6, 2024

Beyond Refusal: Scrubbing Hazards from Open-Source Models

Kyle Gabriel Reynoso, Ivan Enclonar, Lexley Maree Villasis · Team Whitedoor Research PH

Submitted to AI and Democracy Hackathon: Demonstrating the Risks. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Beyond Refusal: Scrubbing Hazards from Open-Source Models

Share

Models trained on the recently published Weapons of Mass Destruction Proxy (WMDP) benchmark show potential robustness in safety due to being trained to forget hazardous information while retaining essential facts instead of refusing to answer. We aim to red-team this approach by answering the following questions on the generalizability of the training approach and its practical scope (see A2).

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Results seem to support claim about unlearning. There are also other approaches to prevent misuse from open-models. https://arxiv.org/abs/2211.14946
    Alternative to unlearning: https://arxiv.org/abs/2404.12699

    When the paper refers to fine-tuning it seems to refer to the unlearning fine-tuning of harmful knowledge. Maybe the wording could sometimes be a bit more clear on this.

    For the refusal vector there was this recent post:
    https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in-llms-is-mediated-by-a-single-direction

    I also am working on a post on refusal vectors in agentic systems.

  2. Interesting experiments, I liked the approach of applying more adversarial pressure to unlearning techniques. Would be interesting to run similar experiments on other unlearning techniques

  3. Great project! I think it’s really important to red-team AI safety methods and your project is a great stab at red-teaming unlearning!

Cite this project

@misc{reynoso2024beyond,
  title = {{Beyond Refusal: Scrubbing Hazards from Open-Source Models}},
  author = {Kyle Gabriel Reynoso and Ivan Enclonar and Lexley Maree Villasis},
  year = {2024},
  month = may,
  note = {Submitted to AI and Democracy Hackathon: Demonstrating the Risks, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/beyond-refusal-scrubbing-hazards-from-open-source-models}},
  url = {https://apartresearch.com/sprints/projects/beyond-refusal-scrubbing-hazards-from-open-source-models}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026