BBLLM
Joey SKAF, Mickaël Boillaud, Thaïs Distinguin · Team BBLLM team
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project focuses on enhancing feature interpretability in large language models (LLMs) by visualizing relationships between latent features. Using an interactive graph-based representation, the tool connects co-activated features for specific prompts, enabling intuitive exploration of feature clusters. Deployed as a web application for Llama-3-70B and Llama-3-8B, it provides insights into the organization of latent features and their roles in decision-making processes.
Reviews
This project develops a visualisation tool for language model SAE latents. Visualisation is an important and underexplored area in interpretability, so it's cool to see this. The visualisation is a graph, where features are connected to one another if they co-occur sufficiently frequently.
The tool is interesting but I'd really like to see some example of the kind of application it might be used in, or an interesting insight (even something very minor) that the authors obtained from using the tool.
This work presents a way to visualise SAE latents that frequently activate together. With some additional time, it'd be cool to see some insights gained from this kind of technique! It'd be especially cool if there was some insight that wasn't easily uncovered via something like UMAP or PCA over the dictionary vectors.
Very interesting visualisation tool! It would have been great to see a bit more of lit.review and see there is specific valua added where other existing techniques fall short. The fact that is ready for local deployment definately deserves extra points.
Cite this project
@misc{skaf2024bbllm,
title = {{BBLLM}},
author = {Joey SKAF and Mickaël Boillaud and Thaïs Distinguin},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/bbllm}},
url = {https://apartresearch.com/sprints/projects/bbllm}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering …