Looking forward to posterity: what past information is transferred to the future?
Zmavli Caimle · Team Nonkansa
Submitted to Computational Mechanics Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
I used mechanistic interpretability techniques to try to see what information the provided Random Randox XOR transformer looks at when making predictions by examining its attention heads manually. I find that earlier layers pay more attention to the previous two tokens, which would be necessary for computing XOR, than the later layers. This finding seems to contradict the finding that more complex computation typically occurs in later layers.

Reviews
No public critique yet.
Cite this project
@misc{caimle2024looking,
title = {{Looking forward to posterity: what past information is transferred to the future?}},
author = {Zmavli Caimle},
year = {2024},
month = jun,
note = {Submitted to Computational Mechanics Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/looking-forward-to-posterity-what-past-information-is-transferred-to-the-future}},
url = {https://apartresearch.com/sprints/projects/looking-forward-to-posterity-what-past-information-is-transferred-to-the-future}
}More from Computational Mechanics Hackathon
- View project: Unsupervised Recovery of Hidden Markov Models from Transformers with Evolutionary Algorithms
Unsupervised Recovery of Hidden Markov Models from Transformers with Evolutionary Algorithms
Stargazy Pie
Prior work finds that transformer neural networks trained to mimic the output of a Hidden Markov Model (HMM) embed the optimal Bayesian beliefs for the HMM's current state in their residual stream, which can be …
- View project: Investigating the Effect of Model Capacity Constraints on Belief State Representations
Investigating the Effect of Model Capacity Constraints on Belief State Representations
Studying Generalization and Abstraction
Computational mechanics provides a formal framework for understanding the concepts needed to perform optimal prediction. Abstraction and generalization seem core to the function of intelligent systems, but are not yet …
- View project: Belief State Representations in Transformer Models on Nonergodic Data
Belief State Representations in Transformer Models on Nonergodic Data
Nerds
We extend research that finds representations of belief spaces in the activations of small transformer models, by discovering that the phenomenon also occurs when the training data stems from Hidden Markov Models whose …