Latent Knowledge Analysis via Feature-Based Causal Tracing
Letlotlo · Team Individual submission
Submitted to Women in AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project explores how factual knowledge is stored in large language models using Goodfire’s Ember API. By identifying and manipulating internal features related to specific facts, it shows how facts are encoded and how model behavior changes when those features are amplified or erased.
Reviews
Good work this weekend! The paper builds upon existing literature (research carried out by Burns et al) - I would recommend reading more widely around the topic to get a better sense of the space. The study identifies general risks posed by LLMs, moving forwards, I would suggest delving into these risks more deeply. The paper addresses key limitations - the focus on one example fact, for example - I would strongly encouraging broadening the scope of the research so that results are more broadly applicable.
Cite this project
@misc{letlotlo2025latent,
title = {{Latent Knowledge Analysis via Feature-Based Causal Tracing}},
author = {Letlotlo},
year = {2025},
month = mar,
note = {Submitted to Women in AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/latent-knowledge-analysis-via-feature-based-causal-tracing}},
url = {https://apartresearch.com/sprints/projects/latent-knowledge-analysis-via-feature-based-causal-tracing}
}More from Women in AI Safety Hackathon
- Education track prizeView project: Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph
One-liner: Morph is the ultimate operation stack for AI safety education—combining dynamic localization, policy simulations, and ecosystem tools to turn abstract risks into actionable, culturally relevant solutions for …
- Mechanistic Interpretability PrizeView project: Red-teaming with Mech-Interpretability
Red-teaming with Mech-Interpretability
Red teaming large language models (LLMs) is crucial for identifying vulnerabilities before deployment, yet systematically creating effective adversarial prompts remains challenging. This project introduces a novel …
- Social Sciences track prizeView project: Detecting Malicious AI Agents Through Simulated Interactions
Detecting Malicious AI Agents Through Simulated Interactions
SafeAIGuard
This research investigates malicious AI Assistants’ manipulative traits and whether the behaviours of malicious AI Assistants can be detected when interacting with human-like simulated users in various decision-making …