Unsupervised Recovery of Hidden Markov Models from Transformers with Evolutionary Algorithms
Dylan Bowman, Colin Lu · Team Stargazy Pie
Submitted to Computational Mechanics Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Prior work finds that transformer neural networks trained to mimic the output of a Hidden Markov Model (HMM) embed the optimal Bayesian beliefs for the HMM's current state in their residual stream, which can be recovered via linear regression. In this work, we aim to address the problem of extracting information about the underlying HMM using the residual stream, without needing to know the MSP already. To do so, we use the R^2 of the linear regression as a reward signal for evolutionary algorithms, which are deployed to search for the parameters that generated the source HMM. We find that for toy scenarios where the HMM is generated by a small set of latent variables, the $R^2$ reward signal is remarkably smooth and the evolutionary algorithms succeed in approximately recovering the original HMM. We believe this work constitutes a promising first step towards the ultimate goal of extracting information about the underlying predictive and generative structure of sequences, by analyzing transformers in the wild.
Reviews
No public critique yet.
Cite this project
@misc{bowman2024unsupervised,
title = {{Unsupervised Recovery of Hidden Markov Models from Transformers with Evolutionary Algorithms}},
author = {Dylan Bowman and Colin Lu},
year = {2024},
month = jun,
note = {Submitted to Computational Mechanics Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/unsupervised-recovery-of-hidden-markov-models-from-transformers-with-evolutionary-algorithms}},
url = {https://apartresearch.com/sprints/projects/unsupervised-recovery-of-hidden-markov-models-from-transformers-with-evolutionary-algorithms}
}More from Computational Mechanics Hackathon
- View project: Looking forward to posterity: what past information is transferred to the future?
Looking forward to posterity: what past information is transferred to the future?
Nonkansa
I used mechanistic interpretability techniques to try to see what information the provided Random Randox XOR transformer looks at when making predictions by examining its attention heads manually. I find that earlier …
- View project: Investigating the Effect of Model Capacity Constraints on Belief State Representations
Investigating the Effect of Model Capacity Constraints on Belief State Representations
Studying Generalization and Abstraction
Computational mechanics provides a formal framework for understanding the concepts needed to perform optimal prediction. Abstraction and generalization seem core to the function of intelligent systems, but are not yet …
- View project: Belief State Representations in Transformer Models on Nonergodic Data
Belief State Representations in Transformer Models on Nonergodic Data
Nerds
We extend research that finds representations of belief spaces in the activations of small transformer models, by discovering that the phenomenon also occurs when the training data stems from Hidden Markov Models whose …