Investigating the Effect of Model Capacity Constraints on Belief State Representations
Ari Brill, Chu Chen · Team Studying Generalization and Abstraction
Submitted to Computational Mechanics Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Computational mechanics provides a formal framework for understanding the concepts needed to perform optimal prediction. Abstraction and generalization seem core to the function of intelligent systems, but are not yet well understood. Computational mechanics may present a promising approach to studying these capabilities. As a preliminary exploration, we examine the effect of weight decay on the fractal structure of belief state representations in a transformer’s residual stream. We find that models trained with increasing weight decay coefficients learn increasingly coarse-grained belief state representations.
Reviews
No public critique yet.
Cite this project
@misc{brill2024investigating,
title = {{Investigating the Effect of Model Capacity Constraints on Belief State Representations}},
author = {Ari Brill and Chu Chen},
year = {2024},
month = jun,
note = {Submitted to Computational Mechanics Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/investigating-the-effect-of-model-capacity-constraints-on-belief-state-representations}},
url = {https://apartresearch.com/sprints/projects/investigating-the-effect-of-model-capacity-constraints-on-belief-state-representations}
}More from Computational Mechanics Hackathon
- View project: Unsupervised Recovery of Hidden Markov Models from Transformers with Evolutionary Algorithms
Unsupervised Recovery of Hidden Markov Models from Transformers with Evolutionary Algorithms
Stargazy Pie
Prior work finds that transformer neural networks trained to mimic the output of a Hidden Markov Model (HMM) embed the optimal Bayesian beliefs for the HMM's current state in their residual stream, which can be …
- View project: Looking forward to posterity: what past information is transferred to the future?
Looking forward to posterity: what past information is transferred to the future?
Nonkansa
I used mechanistic interpretability techniques to try to see what information the provided Random Randox XOR transformer looks at when making predictions by examining its attention heads manually. I find that earlier …
- View project: Belief State Representations in Transformer Models on Nonergodic Data
Belief State Representations in Transformer Models on Nonergodic Data
Nerds
We extend research that finds representations of belief spaces in the activations of small transformer models, by discovering that the phenomenon also occurs when the training data stems from Hidden Markov Models whose …