Skip to content
Sprint projectMar 10, 2025

Debugging Language Models with SAEs

Wen Xing · Team SAE Mechanic

Submitted to Women in AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Debugging Language Models with SAEs

Code (opens in new tab)
Share

This report investigates an intriguing failure mode in the Llama-3.1-8B-Instruct model: its inconsistent ability to count letters depending on letter case and grammatical structure. While the model correctly answers "How many Rs are in BERRY?", it struggles with "How many rs are in berry?", suggesting that uppercase and lowercase queries activate entirely different cognitive pathways. Through Sparse Autoencoder (SAE) analysis, feature activation patterns reveal that uppercase queries trigger letter-counting features, while lowercase queries instead activate uncertainty-related neurons. Feature steering experiments show that simply amplifying counting neurons does not lead to correct behavior. Further analysis identifies tokenization effects as another important factor: different ways of breaking very similar sentences into tokens influence the model’s response. Additionally, grammatical structure plays a role, with "is" phrasing yielding better results than "are."

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Interesting research direction - well done! This paper presents an innovative approach to an interesting issue - how capitalisations and grammar in prompts can impact results. I would suggest carrying out a thorough literature review to understand the problem space in a lot of detail. I'd also think carefully about AI safety risks and what token inconsistencies could mean. This is an interesting research direction and could be made stronger by engaging with the literature and potential impacts on the safety space.

Cite this project

@misc{xing2025debugging,
  title = {{Debugging Language Models with SAEs}},
  author = {Wen Xing},
  year = {2025},
  month = mar,
  note = {Submitted to Women in AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/debugging-language-models-with-saes}},
  url = {https://apartresearch.com/sprints/projects/debugging-language-models-with-saes}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026