Skip to content
Sprint projectMar 10, 2025

Mechanistic Interpretability Track: Neuronal Pathway Coverage

Garance Colomer, Luc Chan, Sacha Lahlou, Leina Corporan Miath · Team 42-Shot

Submitted to Women in AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Mechanistic Interpretability Track: Neuronal Pathway Coverage

Presentation

Presentation: Mechanistic Interpretability Track: Neuronal Pathway Coverage

Code (opens in new tab)
Share

Our study explores mechanistic interpretability by analyzing how Llama 3.3 70B classifies political content. We first infer user political alignment (Biden, Trump, or Neutral) based on tweets, descriptions, and locations. Then, we extract the most activated features from Biden- and Trump-aligned datasets, ranking them based on stability and relevance. Using these features, we reclassify users by prompting the model to rely only on them. Finally, we compare the new classifications with the initial ones, assessing neural pathway overlap and classification consistency through accuracy metrics and visualization of activation patterns.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Studying the representation of different groups of people (e.g. representation of different political affiliations) in LLMs is a very interesting question! A more detailed methods section (e.g. show the prompts fed to the model at each step; a figure with an example of a tweet, the feature activations, and then the classification) would help make some of the specifics easier to follow.

Cite this project

@misc{colomer2025mechanistic,
  title = {{Mechanistic Interpretability Track: Neuronal Pathway Coverage}},
  author = {Garance Colomer and Luc Chan and Sacha Lahlou and Leina Corporan Miath},
  year = {2025},
  month = mar,
  note = {Submitted to Women in AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/mechanistic-interpretability-track-neuronal-pathway-coverage}},
  url = {https://apartresearch.com/sprints/projects/mechanistic-interpretability-track-neuronal-pathway-coverage}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026