Debugging Language Models with SAEs
Wen Xing · Team SAE Mechanic
Submitted to Women in AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This report investigates an intriguing failure mode in the Llama-3.1-8B-Instruct model: its inconsistent ability to count letters depending on letter case and grammatical structure. While the model correctly answers "How many Rs are in BERRY?", it struggles with "How many rs are in berry?", suggesting that uppercase and lowercase queries activate entirely different cognitive pathways. Through Sparse Autoencoder (SAE) analysis, feature activation patterns reveal that uppercase queries trigger letter-counting features, while lowercase queries instead activate uncertainty-related neurons. Feature steering experiments show that simply amplifying counting neurons does not lead to correct behavior. Further analysis identifies tokenization effects as another important factor: different ways of breaking very similar sentences into tokens influence the model’s response. Additionally, grammatical structure plays a role, with "is" phrasing yielding better results than "are."
Reviews
Interesting research direction - well done! This paper presents an innovative approach to an interesting issue - how capitalisations and grammar in prompts can impact results. I would suggest carrying out a thorough literature review to understand the problem space in a lot of detail. I'd also think carefully about AI safety risks and what token inconsistencies could mean. This is an interesting research direction and could be made stronger by engaging with the literature and potential impacts on the safety space.
Cite this project
@misc{xing2025debugging,
title = {{Debugging Language Models with SAEs}},
author = {Wen Xing},
year = {2025},
month = mar,
note = {Submitted to Women in AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/debugging-language-models-with-saes}},
url = {https://apartresearch.com/sprints/projects/debugging-language-models-with-saes}
}More from Women in AI Safety Hackathon
- Education track prizeView project: Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph
One-liner: Morph is the ultimate operation stack for AI safety education—combining dynamic localization, policy simulations, and ecosystem tools to turn abstract risks into actionable, culturally relevant solutions for …
- Mechanistic Interpretability PrizeView project: Red-teaming with Mech-Interpretability
Red-teaming with Mech-Interpretability
Red teaming large language models (LLMs) is crucial for identifying vulnerabilities before deployment, yet systematically creating effective adversarial prompts remains challenging. This project introduces a novel …
- Social Sciences track prizeView project: Detecting Malicious AI Agents Through Simulated Interactions
Detecting Malicious AI Agents Through Simulated Interactions
SafeAIGuard
This research investigates malicious AI Assistants’ manipulative traits and whether the behaviours of malicious AI Assistants can be detected when interacting with human-like simulated users in various decision-making …