The next jump
Mihir Sahasrabudhe · Team MS
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
As we plan our next major training run, we look at the pressures: increased capability, reduced monitorability, and reward hacking in novel ways. We assess the current literature around these issues and underline the need of stronger monitors.
Reviews
This project is a thoughtful and well-written piece. I liked how it is based on the actual reports (the OpenAI-HF and artifactory incidents, and the Mythos monitor result showing how the CoT made the monitor worse). I also liked the way "strength of the monitor" is framed as a hyperparameter with both a strict and loose tradeoff - this is a clear way to think about it. One limitation I think about is the empirical part being small (47 R-Judge records, one model, four conditions that all made identical decisions), so it mostly illustrates the argument rather than testing it. Overall I think this is a strong, clearly reasoned position paper, and to me the next step would be turning this argument into a bigger test.
A small experiment, unusually well run. The configuration was frozen before the run — model, seed, sampling, budget caps, provider pinned, the call schedule fixed in advance — and the development and evaluation splits share no records and no task families, with development closed out first. Independent reconstruction from the raw logs reproduces the reported table exactly, and more tellingly, the same cases are missed in every condition rather than the counts merely matching. A failed call is recorded with its cost unresolved instead of quietly written off.
The negative result is the valuable part and is not oversold: extra context, an independent critique step and simple repetition all changed nothing, because every layer made the same error of treating authorization as sufficient.
Two things. Almost none of this methodology appears in the paper, which gives away work that is part of the contribution. And the empirical footprint is a single page — one model, one benchmark, a modest set of records, a borrowed technique. A second model would turn a result about this particular test into a claim about the blind spot itself.
Read full reviewShow less
Cite this project
@misc{sahasrabudhe2026next,
title = {{The next jump}},
author = {Mihir Sahasrabudhe},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-next-jump-4h36}},
url = {https://apartresearch.com/sprints/projects/the-next-jump-4h36}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …