Skip to content
Sprint projectJul 28, 2025Monoid di Safety Hub

RL vs Active Inference with respect to Reward Hacking

Grigoriy Erdyakov, Alexey Biriulin, Arseniy Varaksin · Team Arseniy

Submitted to AI Safety x Physics Grand Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: RL vs Active Inference with respect to Reward Hacking

Share

We aim to apply Reinforcement Learning (RL) and Active Inference methods to implement intelligent agents in various environments and compare their behavioral characteristics, particularly failure modes similar to reward hacking, using relevant detectors. We hypothesize that agents based on Active Inference may exhibit safer behavior.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How rigorous is your physics methodology and how feasible is your approach? Is your theoretical framework sound and your empirical work well-designed? Can your proposed methods be implemented and validated?

How clearly does your work address important AI safety challenges? What is the potential impact on ensuring beneficial AI development? Does your approach offer meaningful insights for AI alignment research?

How novel and creative is your approach to bridging physics and AI safety? Do you introduce new theoretical connections or methodological innovations? What makes your work distinct from existing research?

  1. AI Safety Relevance: Reward hacking is an important issue in AI Safety, and finding RL algorithms which are biased against it is an important problem to study. While I am curious to see how active inference agents would perform on different reward hacking measures, I am personally skeptical that positive results here would lead towards algorithms with a sufficiently low alignment tax. I say this with low-medium confidence, as I am not heavily acquainted, much less an expert on FEP, and am mostly deferring to the bearish picture painted by Beren in this post: https://www.beren.io/2024-07-27-A-Retrospective-on-Active-Inference/.

    Innovation & Originality: Proponents of Active Inference will often highlight it's exploration advantages when comparing it to traditional RL methods. So in that regard it is well motivated.

    Technical Rigor & Feasibility: Despite a few passages being unclear, the authors managed to ask a reasonable set of questions, had an informed engagement with the literature and laid out a reasonable research plan.

    Read full reviewShow less
  2. No concrete results, no code and no physics analogy.

  3. The submission proposes to study active inference as an alternative to RL as a model for AI agents. Active inference is an influential theory in neuroscience that has not been fully explored in AI safety, and the field might benefit from having a better understanding of the circumstances in which it can usefully describe AI systems. I found it thought-provoking how the proposal compared active inference and RL through the lens of alignment-relevant failure modes such as reward hacking and wireheading. The proposal would have been improved by including a more extensive overview and literature review of active inference, and/or by providing a starting point for further research by discussing in detail one or more specific scenarios in which active inference and RL differ in an interesting way.

Cite this project

@misc{erdyakov2025rl,
  title = {{RL vs Active Inference with respect to Reward Hacking}},
  author = {Grigoriy Erdyakov and Alexey Biriulin and Arseniy Varaksin},
  year = {2025},
  month = jul,
  note = {Submitted to AI Safety x Physics Grand Challenge, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/rl-vs-active-inference-with-respect-to-reward-hacking-8h3s}},
  url = {https://apartresearch.com/sprints/projects/rl-vs-active-inference-with-respect-to-reward-hacking-8h3s}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026