Skip to content
Sprint projectNov 25, 2024

Analyzing Dataset Bias with SAEs

Nick Jiang, Joseph Tey · Team Big gamers

Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Analyzing Dataset Bias with SAEs

Share

We use SAEs to study biases in datasets.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project contribute to the field of mechanistic interpretability? Does it provide new insights into understanding or steering AI model behavior? How well does it move us towards reprogramming AI models? How original and innovative is the approach?

How important is the contribution to advancing the field of AI safety? Do we expect the results to generalize beyond the specific case(s) presented in the submission? Does the approach introduce new safety mechanisms or enhance existing ones in innovative ways?

How well is the project executed from a technical standpoint?, Is the code well-structured, documented, and reproducible?, How effectively does it utilize Goodfire's SDK/API and other provided resources?, How clearly and effectively is the research presented in the paper?, Quality of visualizations and demos (if applicable), Clarity of methodology explanation and results interpretation

  1. This paper is about identifying issues in the quality of pre-training datasets. They look for features that light up for certain classes, such as spam or buggy code, using contrastive search. They then see if these features are really meaningfully. It maybe should have been causally instead of casually: "causally ties a predicted output with activated features." The basic idea seems to be: We have two classes of text, buggy code vs safe code, now we use contrastive search to find features that separate them. After this step, we try to figure out exactly what those features fire for; potentially, we may find features that shouldn't be there. The imagined success could look like this: we train the model on a dataset in which there is some spurious correlation between buggy code and something else. (For example, imagine a programmer in a company creates a lot of buggy code and also has a habit of using a certain library a lot. A spurious correlation could cause models to flag code including this library.) We may want to remove either this data or the feature. It seems that they haven't or don't explain a process to potentially automatically detect such correlations. For me, it is, however, not clear that there really is an established problem here. Could be good to at least demonstrate or cite issues with spurious correlations in current models unless I missed this.

    Read full reviewShow less
  2. I really like the application idea here. It would be amazing to understand better the downstream implications of some of the biases identified so far on model performance and really work out the threat model being addressed here in detail. This seems worthy of followup work.

  3. I think that using SAE features to find different biases and information about a dataset seems like a worthwhile direction. However, I'm struggling to understand what each of the things you did are trying to achieve. I think that this comes from using too many methods that make understanding the end goal of the final result blurry.

Cite this project

@misc{jiang2024analyzing,
  title = {{Analyzing Dataset Bias with SAEs}},
  author = {Nick Jiang and Joseph Tey},
  year = {2024},
  month = nov,
  note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/analyzing-dataset-bias-with-saes}},
  url = {https://apartresearch.com/sprints/projects/analyzing-dataset-bias-with-saes}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026