Kobayashi Maru: a pre-registered dose–response study of cheating spillover from impossible tasks to the solvable ones beside them
Ebin Babu Thomas · Team Kobayashi Maru
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Incident investigations blamed the July 2026 OpenAI/Hugging Face breach partly on ExploitGym's 30–40% impossible tasks creating cheating pressure, but nobody had varied that fraction to test it. I did, over 8,959 agent runs. Every batch held the same ten solvable Python tasks plus a varying number of impossible ones, and I measured cheating only on the solvable tasks. Two of the four models that took the bait cheated more as the batch filled, the largest going from 0% to 30%. The carrier is the agent's own notes being replayed back to it: withhold them and the spillover disappears entirely. But nothing the agents submitted was wrong, so what spreads is reconnaissance, not an exploit.
Reviews
Strong experimental execution. The result that replayed agent notes can carry the behavior forward is interesting because it points to a concrete thing eval designers can monitor. My main question is how consequential the effect is outside this setup. I wld next test this without a planted answer file, on harder tasks where accessing the grader can actually change the submitted result, and try selective note-scrubbing rather than removing the whole memory channel.
Interesting idea, to the best of my knowledge a novel direction and one which it seems important to understand. The conclusions regarding agent memory and reconnaissance habits leading to cheating are potentially impactful. They were nonetheless not really highlighted in the abstract or introduction.
I would have liked to see more consideration of whether merely shuffling the impossible tasks in was sufficient - presumably the agents are still going to do the eval serially and it's much more relevant the fraction of impossible tasks they've seen so far than in the whole trajectory?
Including the exact numbers in the abstract was probably unnecessary and made it harder to read. Additionally it does not present what I think are the most important take-aways from the project.
Much of the text reads as very AI-generated which makes it hard to believe the claims are quite correct.
I think this paper asks an interesting question about whether impossible tasks encourage agents to break rules on nearby tasks they could solve. Measuring behavior on the solvable tasks is a useful contribution, and the reported negative results help show where the effect did not appear.
However, adding impossible tasks also makes the batches longer, so the experiment does not clearly separate those explanations. I would compare equally long batches with and without impossible tasks before making the stronger causal claim. I also found the paper harder to review than the question warranted. Dense statistical terminology, especially in the abstract, repeated claims, and secondary analyses makes it difficult to follow the main experiment or for others to build on the work. A clearer account separating the findings from their interpretation without dense terminology would make the contribution easier to judge.
Read full reviewShow less
Cite this project
@misc{thomas2026kobayashi,
title = {{Kobayashi Maru: a pre-registered dose–response study of cheating spillover from impossible tasks to the solvable ones beside them}},
author = {Ebin Babu Thomas},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/kobayashi-maru-a-preregistered-doseresponse-study-of-cheating-spillover-from-impossible-tasks-to-the-solvable-ones-beside-them-kwzr}},
url = {https://apartresearch.com/sprints/projects/kobayashi-maru-a-preregistered-doseresponse-study-of-cheating-spillover-from-impossible-tasks-to-the-solvable-ones-beside-them-kwzr}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …