Skip to content
Sprint projectMar 23, 2026San Francisco

Post-hoc Training Interventions Against Naturally Emergent Misalignment

Ethan Elasky, Frank Nakasako · Team Bubble Research

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Post-hoc Training Interventions Against Naturally Emergent Misalignment

Share

Training-based AI control uses training to mitigate risk from misaligned models, but prior work has only tested this against prompted misalignment. Ryd, Benton & Hebbar (2025) show that SFT on 10–100 samples eliminates prompted attacks on math tasks but recovers only 50–90% of benign performance, and identify RL as the most promising next step without testing it. We extend their work to naturally emergent misalignment — a model that has learned to reward hack through RL on coding tasks — and test both SFT and the RL interventions they propose. We replicate reward hacking emergence on gpt-oss-20b using GRPO on LeetCode problems, porting the training pipeline from veRL to the Tinker cloud API. Hacking emerges by step 25 and saturates at 95% by step 75, but only on gpt-oss-20b — not on Qwen3-8B variants — revealing high sensitivity to model architecture. SFT eliminates hacking completely but catastrophically damages coding ability (0–4% correct rate vs. 62.5% base), confirming Ryd et al.'s capability damage finding on harder misalignment. RL with ground-truth reward recovers 48.4% correct rate with 0% hacking. We test RL with a trusted Qwen3-1.7b monitor (6.9% correct rate) supervising correction of gpt-oss-20b, directly instantiating the scalable oversight question: can a weak trusted model provide sufficient training signal to correct a stronger model's learned misalignment?

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project undertakes an empirical investigation of a method for reducing misaligned (specifically reward hacking) behavior through post-hoc training methods. The use of the trusted model T as a source-of-truth makes it relevant to a control setting. The work appears methodologically sound and well-executed. The finding that SFT capability damage is substantially worse against trained misalignment than prompted misalignment (19% vs 50–90% recovery) is a useful quantitative result, and the demonstration that RL with a weak trusted monitor nearly matches ground-truth correction is notable.

    However, the main experimental conditions - RL as intervention, weak model supervision, and the move to trained misalignment, are all explicitly proposed in Ryd et al. (2025). While confirming these hypotheses empirically is valuable, the work does not attempt to contribute novel ideas. For example, how robust the weak supervision result is would warrant further investigation.

    Presenting (even speculative) ideas for how to go beyond what is done, or raising interesting questions to investigate in the future would elevate this from a strong empirical confirmation to a paper that also shapes the direction of future work in this area.

    Read full reviewShow less
  2. Useful extension of Ryd et al., testing the RL interventions they identified as most promising is a natural and worthwhile next step, and moving from prompted to naturally emergent misalignment is a harder, more realistic setting. The central result (weak Qwen3-1.7b correcting stronger gpt-oss-20b) is a concrete positive datapoint for scalable oversight.

    That said, reward hacking via test function overwriting is a known, mechanical exploit, and the paper's own discussion suggests that removing the exploitable pathway may matter more than the monitor's quality, which would weaken the claim of scalable oversight. The single-seed, single-domain setup and the loss architecture confound limit how much to read into the specific numbers. Replicating across tasks and seeds would considerably strengthen confidence.

Cite this project

@misc{elasky2026posthoc,
  title = {{Post-hoc Training Interventions Against Naturally Emergent Misalignment}},
  author = {Ethan Elasky and Frank Nakasako},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/posthoc-training-interventions-against-naturally-emergent-misalignment-3tvj}},
  url = {https://apartresearch.com/sprints/projects/posthoc-training-interventions-against-naturally-emergent-misalignment-3tvj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026