Skip to content
Sprint projectNov 24, 2024

Recovering Goodfire's SAE feature vectors from their API

Lovkush Agarwal · Team Lovkush

Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Recovering Goodfire's SAE feature vectors from their API

Code (opens in new tab)
Share

In this project, we carry out an early trial to see whether Goodfire’s SAE feature vectors can be recovered using the information available from their API. The strategy tried is: pick a feature of interest, construct a contrastive dataset using Goodfire’s API, then use TransformerLens to get a steering vector for the contrastive dataset, by simply calculating the average difference in the activations in each pair.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project contribute to the field of mechanistic interpretability? Does it provide new insights into understanding or steering AI model behavior? How well does it move us towards reprogramming AI models? How original and innovative is the approach?

How important is the contribution to advancing the field of AI safety? Do we expect the results to generalize beyond the specific case(s) presented in the submission? Does the approach introduce new safety mechanisms or enhance existing ones in innovative ways?

How well is the project executed from a technical standpoint?, Is the code well-structured, documented, and reproducible?, How effectively does it utilize Goodfire's SDK/API and other provided resources?, How clearly and effectively is the research presented in the paper?, Quality of visualizations and demos (if applicable), Clarity of methodology explanation and results interpretation

  1. Similar to papers such as Recovering the Input and Output Embedding of OpenAI Models (https://arxiv.org/pdf/2403.06634), this paper seeks to recover the SAE vectors from Goodfire's API. The authors admit that they failed at their objective but generally conclude that it should be possible. The method they propose seems reasonable, but I am confused about a key detail: the steering vectors are typically applied to the residual stream, though it’s not entirely clear at which position. The authors want to cache output activations, though those might be in the token space after applying the output embedding? It seems that their method doesn't quite work for some reason. "The steering vector is the average of all these differences, across n and across the dataset." <- This statement might be too strong; it depends on a lot of factors, such as where the steering vector is applied, which datasets are used, and how contrastive pairs are generated.

    Read full reviewShow less
  2. Interesting study direction, which seems to be very worthwhile for further study. I think that the example that you provide does show some success in this idea. I’m mainly not sure if it makes sense to find the activations in residual stream for features which are monosemantic from an SAE, as the reason for SAEs is that the residual stream itself is polysemantic.

  3. I was excited to see someone try to do this -- I think how easy it is to reconstruct things like the dictionary vectors has important practical implications. This work would be improved by providing more motivation for the technique of choice, perhaps providing an example or figure that demonstrates why it should work.

  4. There seems to be existing very similar work and the results from this trial (as acknoledged from the author) is not very succesful. Overall it is an interesting approach that could benefit from more lit.review and further exploration.

Cite this project

@misc{agarwal2024recovering,
  title = {{Recovering Goodfire's SAE feature vectors from their API}},
  author = {Lovkush Agarwal},
  year = {2024},
  month = nov,
  note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/recovering-goodfire-s-sae-feature-vectors-from-their-api}},
  url = {https://apartresearch.com/sprints/projects/recovering-goodfire-s-sae-feature-vectors-from-their-api}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026