Attention Pattern Based Information Flow Visualization Tool
Sofia Maria Lo Cicero Vaina, Anastasiia Ivanova, Liubov Yaronskaya · Team babushka’s
Submitted to Women in AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Understanding information flow in transformer-based language models is crucial for mechanistic interpretability. We introduce a visualization tool that extracts and represents attention patterns across model components, revealing how tokens influence each other during processing. Our tool automatically identifies and color-codes functional attention head types based on established taxonomies from recent research on indirect object identification (Wang et al., 2022), factual recall (Chughtai et al., 2024), and factual association retrieval (Geva et al., 2023). This interactive approach enables researchers to trace information propagation through transformer architectures, providing deeper insights into how these models implement reasoning and knowledge retrieval capabilities.
Reviews
This paper presents an interactive visualisation tool for attention patterns in transformer models. The authors demonstrate good understanding of the relevant literature and make circuit discovery more accessible and interactive. The tool has limitations in visualising MLP contributions (as highlighted by the authors) and may risk overinterpreting attention patterns, but it's still a really cool contribution!
Very cool project directly addressing a problem area within mechanistic interpretability: visualizing information flow in transformers. By automating the process of identifying and color-coding attention head types, the tool lowers the barrier to entry for researchers and practitioners. It builds directly on established taxonomies from key papers in the field , demonstrating a strong understanding of the existing literature. The methodology is clearly explained and the application of the tool to known circuits provides strong validation. Well done! I would have liked to see a bit more discussion of safety implications, as it is currently somewhat high level - examples would be useful here, and perhaps some mention of computational cost. Expansions beyond the ones identified could also include incorporating methods for identifying polysemanticity in attention heads, which might help uncover more nuanced information flow..
Read full reviewShow less
Cite this project
@misc{vaina2025attention,
title = {{Attention Pattern Based Information Flow Visualization Tool}},
author = {Sofia Maria Lo Cicero Vaina and Anastasiia Ivanova and Liubov Yaronskaya},
year = {2025},
month = mar,
note = {Submitted to Women in AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/attention-pattern-based-information-flow-visualization-tool-e19d}},
url = {https://apartresearch.com/sprints/projects/attention-pattern-based-information-flow-visualization-tool-e19d}
}More from Women in AI Safety Hackathon
- Education track prizeView project: Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph
One-liner: Morph is the ultimate operation stack for AI safety education—combining dynamic localization, policy simulations, and ecosystem tools to turn abstract risks into actionable, culturally relevant solutions for …
- Mechanistic Interpretability PrizeView project: Red-teaming with Mech-Interpretability
Red-teaming with Mech-Interpretability
Red teaming large language models (LLMs) is crucial for identifying vulnerabilities before deployment, yet systematically creating effective adversarial prompts remains challenging. This project introduces a novel …
- Social Sciences track prizeView project: Detecting Malicious AI Agents Through Simulated Interactions
Detecting Malicious AI Agents Through Simulated Interactions
SafeAIGuard
This research investigates malicious AI Assistants’ manipulative traits and whether the behaviours of malicious AI Assistants can be detected when interacting with human-like simulated users in various decision-making …