Searching for Universality and Equivariance in LLMs using Sparse Autoencoder Found Features
Meruyert Alimaganbetova, Jason Zeng · Team Meru & Jason
Submitted to Women in AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
The project investigates how neuron features with properties of universality and equivariance affect the controllability and safety of large language models, finding that behaviors supported by redundant features are more resistant to manipulation than those governed by singular features.
Reviews
I like this research direction a lot and think there's signal to explore re: redundant feature sets bolstering safety-relevant behaviors vs. singular features remaining more vulnerable. It would have been a more robust analysis if the team had been able to generalize across a larger dataset of prompt examples, and also more clearly defined linguistically a formula for "equivariance" or "singular"/"redundant" features (e.g., based on existing literature).
Cite this project
@misc{alimaganbetova2025searching,
title = {{Searching for Universality and Equivariance in LLMs using Sparse Autoencoder Found Features}},
author = {Meruyert Alimaganbetova and Jason Zeng},
year = {2025},
month = mar,
note = {Submitted to Women in AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/searching-for-universality-and-equivariance-in-llms-using-sparse-autoencoder-found-features}},
url = {https://apartresearch.com/sprints/projects/searching-for-universality-and-equivariance-in-llms-using-sparse-autoencoder-found-features}
}More from Women in AI Safety Hackathon
- Education track prizeView project: Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph
One-liner: Morph is the ultimate operation stack for AI safety education—combining dynamic localization, policy simulations, and ecosystem tools to turn abstract risks into actionable, culturally relevant solutions for …
- Mechanistic Interpretability PrizeView project: Red-teaming with Mech-Interpretability
Red-teaming with Mech-Interpretability
Red teaming large language models (LLMs) is crucial for identifying vulnerabilities before deployment, yet systematically creating effective adversarial prompts remains challenging. This project introduces a novel …
- Social Sciences track prizeView project: Detecting Malicious AI Agents Through Simulated Interactions
Detecting Malicious AI Agents Through Simulated Interactions
SafeAIGuard
This research investigates malicious AI Assistants’ manipulative traits and whether the behaviours of malicious AI Assistants can be detected when interacting with human-like simulated users in various decision-making …