Skip to content
Sprint projectApr 27, 2026Olympia, Washington, USA

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Caleb DeLeeuw

Submitted to AIxBio Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Recording (opens in new tab)Code (opens in new tab)More on huggingface.co (opens in new tab)
Share

BioRefusalAudit measures whether a model's refusal is structurally real or just surface-level. Using sparse autoencoder interpretability( both off-the-shelf Gemma Scope SAEs and a biosecurity-specific SAE trained during the hackathon) it computes a divergence score between what a model says and what its internal activations show. Key findings: Gemma 2 never genuinely refuses, it hedges. A single chat-template token takes Gemma 4 from 65 refusals to zero. Both models refuse nothing at 80-token caps. And the refusal circuit fires harder on psilocybin (biologically benign, Schedule I) than on genuinely hazardous biology. None of this is visible to surface evaluation. Runs on a 4 GB consumer GPU. Dozens of different trial runs with various models, contexts, and controls, showing consistent results. Code and data open yet secured under Hippocratic License 3.0 with select modules highlighting its usefulness for biosecurity and AI safety research.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. General notes

    • Interp techniques to see whether refusals are robust or easily bypassible seems broadly useful. The methods seem somewhat shaky – the feature catalog is populated purely based on statistical selection (unclear whether these distinctions are meaningful), contrastive SAE seems to not have achieved the seperation it was optimized for, etc. Difficult to know what conclusion to draw from the specific technical approach implemented here.
    • The finding about the Schedule I pschyadelic probe is the most interesting, suggesting some kind of representation of cultural tabooness.
    • Writing is quite LLM-y and therefore a bit sloppy/annoying to read. Very dramatic, lots of "It's not X, it's Y", colons, etc. Less text but with more substance would be preferable. Main findings/takeaways are quite buried. Policy connections and review/understanding of previous literature seems surface level.
    • Would have been nice to see some example prompts or method behind prompt generation.
    • Recent relevant work the author may be interested in: https://securebio.org/biotier/

    Small notes

    • (Intro P1) LAB-Bench is not designed to measure dangerous biology proxies, just general bio knowledge.
    • (Related work, policy framing) Not sure how this work addresses "tiered access" from Yassif and Carter? It's a refusal benchmark. Tiered access is quite a different system.
    Read full reviewShow less
  2. This started out as a pretty compelling research project that clearly described a big problem (i.e. the difference between model behaviour in terms of the chat output vs the underlying activations), and the major findings were laid out well in the Abstract.

    Although the technical scope is a little outside my expertise, I found that the rest of the report was extremely jargon heavy and seemingly reliant on LLMs to generate the text, which made me suspicious about the analysis of the results. Overall this ‘felt like’ a good contribution to interpretability w.r.t biosecurity, but a more technical judge with familiarity of interpretability would have a fairer view.

    Some pointers

    Some explanation in the intro was clunky ‘The question is: when a model refuses, does it can’t, or does it merely won’t right now’ → I think I get what you mean, but a few phrases were pretty confusing.

    The more I read of this document, the more text was either extremely jargon heavy or LLM-flavoured (e.g. the entire Related Work section, and particularly the ‘Policy framing’ section seemed pretty LLM coded). Then in the results and limitations/discussion, whole passages seemed very LLM-y to me, which made me question the appraisal of the results.

    The D metric is pretty important, but not particularly well explained. You formally define it, but that’s not that helpful to non-AI/maths types. Even for a somewhat AI-literate person, I wasn’t left with much intuition about what D is actually measuring, or how novel or valid it is (I imagine something similar has been done for other interpretability efforts?)..

    Read full reviewShow less
  3. I like how sharp the contribution of this methodology is. I would suggest restructuring the framing, perhaps leading with behavioral findings (format gating, schedule I findings, etc.) since it's more robust compared to the metric D centerpiece framing. I also think that the single model family limitation is actually more significant than given light to, the mechanistic claims made here rest only on one model. The dual use consideration is well handle.

Cite this project

@misc{deleeuw2026biorefusalaudit,
  title = {{BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders}},
  author = {Caleb DeLeeuw},
  year = {2026},
  month = apr,
  note = {Submitted to AIxBio Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/biorefusalaudit-auditing-biosecurity-refusal-depth-using-general-and-domainfinetuned-sparse-autoencoders-1fyk}},
  url = {https://apartresearch.com/sprints/projects/biorefusalaudit-auditing-biosecurity-refusal-depth-using-general-and-domainfinetuned-sparse-autoencoders-1fyk}
}
Browse all projects

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026