Skip to content
Sprint projectMar 10, 2025

Searching for Universality and Equivariance in LLMs using Sparse Autoencoder Found Features

Meruyert Alimaganbetova, Jason Zeng · Team Meru & Jason

Submitted to Women in AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Searching for Universality and Equivariance in LLMs using Sparse Autoencoder Found Features

Code (opens in new tab)
Share

The project investigates how neuron features with properties of universality and equivariance affect the controllability and safety of large language models, finding that behaviors supported by redundant features are more resistant to manipulation than those governed by singular features.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. I like this research direction a lot and think there's signal to explore re: redundant feature sets bolstering safety-relevant behaviors vs. singular features remaining more vulnerable. It would have been a more robust analysis if the team had been able to generalize across a larger dataset of prompt examples, and also more clearly defined linguistically a formula for "equivariance" or "singular"/"redundant" features (e.g., based on existing literature).

Cite this project

@misc{alimaganbetova2025searching,
  title = {{Searching for Universality and Equivariance in LLMs using Sparse Autoencoder Found Features}},
  author = {Meruyert Alimaganbetova and Jason Zeng},
  year = {2025},
  month = mar,
  note = {Submitted to Women in AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/searching-for-universality-and-equivariance-in-llms-using-sparse-autoencoder-found-features}},
  url = {https://apartresearch.com/sprints/projects/searching-for-universality-and-equivariance-in-llms-using-sparse-autoencoder-found-features}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026