Belief State Representations in Transformer Models on Nonergodic Data
Junfeng Feng, Wanjie Zhong, Doroteya Stoyanova, Lennart Finke · Team Nerds
Submitted to Computational Mechanics Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We extend research that finds representations of belief spaces in the activations of small transformer models, by discovering that the phenomenon also occurs when the training data stems from Hidden Markov Models whose hidden states do not communicate at all. Our results suggest that Bayesian updating and internal belief state representation also occur when they are not necessary to perform well in the prediction task, providing tentative evidence that large transformers keep a representation of their external world as well.

Reviews
No public critique yet.
Cite this project
@misc{feng2024belief,
title = {{Belief State Representations in Transformer Models on Nonergodic Data}},
author = {Junfeng Feng and Wanjie Zhong and Doroteya Stoyanova and Lennart Finke},
year = {2024},
month = jun,
note = {Submitted to Computational Mechanics Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/belief-state-representations-in-transformer-models-on-nonergodic-data}},
url = {https://apartresearch.com/sprints/projects/belief-state-representations-in-transformer-models-on-nonergodic-data}
}More from Computational Mechanics Hackathon
- View project: Unsupervised Recovery of Hidden Markov Models from Transformers with Evolutionary Algorithms
Unsupervised Recovery of Hidden Markov Models from Transformers with Evolutionary Algorithms
Stargazy Pie
Prior work finds that transformer neural networks trained to mimic the output of a Hidden Markov Model (HMM) embed the optimal Bayesian beliefs for the HMM's current state in their residual stream, which can be …
- View project: Looking forward to posterity: what past information is transferred to the future?
Looking forward to posterity: what past information is transferred to the future?
Nonkansa
I used mechanistic interpretability techniques to try to see what information the provided Random Randox XOR transformer looks at when making predictions by examining its attention heads manually. I find that earlier …
- View project: Investigating the Effect of Model Capacity Constraints on Belief State Representations
Investigating the Effect of Model Capacity Constraints on Belief State Representations
Studying Generalization and Abstraction
Computational mechanics provides a formal framework for understanding the concepts needed to perform optimal prediction. Abstraction and generalization seem core to the function of intelligent systems, but are not yet …