RL vs Active Inference with respect to Reward Hacking
Grigoriy Erdyakov, Alexey Biriulin, Arseniy Varaksin · Team Arseniy
Submitted to AI Safety x Physics Grand Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.
We aim to apply Reinforcement Learning (RL) and Active Inference methods to implement intelligent agents in various environments and compare their behavioral characteristics, particularly failure modes similar to reward hacking, using relevant detectors. We hypothesize that agents based on Active Inference may exhibit safer behavior.
Reviews
AI Safety Relevance: Reward hacking is an important issue in AI Safety, and finding RL algorithms which are biased against it is an important problem to study. While I am curious to see how active inference agents would perform on different reward hacking measures, I am personally skeptical that positive results here would lead towards algorithms with a sufficiently low alignment tax. I say this with low-medium confidence, as I am not heavily acquainted, much less an expert on FEP, and am mostly deferring to the bearish picture painted by Beren in this post: https://www.beren.io/2024-07-27-A-Retrospective-on-Active-Inference/.
Innovation & Originality: Proponents of Active Inference will often highlight it's exploration advantages when comparing it to traditional RL methods. So in that regard it is well motivated.
Technical Rigor & Feasibility: Despite a few passages being unclear, the authors managed to ask a reasonable set of questions, had an informed engagement with the literature and laid out a reasonable research plan.
Read full reviewShow less
No concrete results, no code and no physics analogy.
The submission proposes to study active inference as an alternative to RL as a model for AI agents. Active inference is an influential theory in neuroscience that has not been fully explored in AI safety, and the field might benefit from having a better understanding of the circumstances in which it can usefully describe AI systems. I found it thought-provoking how the proposal compared active inference and RL through the lens of alignment-relevant failure modes such as reward hacking and wireheading. The proposal would have been improved by including a more extensive overview and literature review of active inference, and/or by providing a starting point for further research by discussing in detail one or more specific scenarios in which active inference and RL differ in an interesting way.
Cite this project
@misc{erdyakov2025rl,
title = {{RL vs Active Inference with respect to Reward Hacking}},
author = {Grigoriy Erdyakov and Alexey Biriulin and Arseniy Varaksin},
year = {2025},
month = jul,
note = {Submitted to AI Safety x Physics Grand Challenge, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/rl-vs-active-inference-with-respect-to-reward-hacking-8h3s}},
url = {https://apartresearch.com/sprints/projects/rl-vs-active-inference-with-respect-to-reward-hacking-8h3s}
}More from AI Safety x Physics Grand Challenge
- Prize winnerView project: Constrained Belief Updates and Geometric Structures in Transformer Representations for the RRXOR Process
Constrained Belief Updates and Geometric Structures in Transformer Representations for the RRXOR Process
In this work, we seek to further understand the computational structure formed by transformers trained to predict the next token of a data set. In particular, we extend the work of (Piotrowski et al., 2025) by analyzing …
- Prize winnerView project: Momentum–Point-Perplexity Mechanics in Large Language Models=
Momentum–Point-Perplexity Mechanics in Large Language Models=
Fractal
This work analyzes the hidden states of twenty different open-source transformer language models, ranging from small to medium size and covering five major architectures. The key discovery is that these models show …
- Prize winnerView project: AI agentic system epidemiology
AI agentic system epidemiology
XIA-for-Med
As AI systems scale into decentralized, multi-agent deployments, emergent vulnerabilities challenge our ability to evaluate and manage systemic risks. In this work, we adapt classical epidemiological modeling …