FunnelBench: Effects of Challenge Framing and Difficulty on Model’s Willingness to Cheat
Bradley Simpkins
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Recent incidents with frontier models have shown that models are, at times, willing to conduct malicious actions to obtain answers to difficult sets of evaluations when capabilities are being tested. This presents a need to create a new methodology that seeks to measure and provide guidance in preventing this type of behavior. FunnelBench seeks to showcase a potential proof of concept that can be further developed to measure a model’s willingness to cheat, as defined by UK AISI3, and showcase how the willingness changes depending upon challenge difficulty and challenge framing. Initial runs showcase that the model tested will read and utilize the answer key more often as challenge tasking becomes unsolvable, and showcases the necessity to measure the multi-stage approach to cheating.
Reviews
The funnel is the contribution and Figure 6 is where it pays off. Splitting access, read and use into three separate measurements, then reporting the conversion between them, tells you something the single-number cheating rates from AISI cannot. I have not seen that decomposition reported elsewhere.
I checked your tables against each other and they hold exactly:
1. Appendix A3 means – Easy 12.5%, Hard 0%, Impossible Unsat 37.5%, Impossible Unreachable 82.5%. Every figure in your Results text matches the appendix.
2. Figure 6 reconciles cell by cell with A4 and A5. Neutral 8/10 read and 3 used gives 38%. Raised Stakes 8/10 and 7 gives 88%. Forbidden 5/10 and 1 gives 20%. Time Pressure 10/10 and 8 gives 80%. All four match your stated conversions.
Three tables built from 200 hand-scored runs, reconciling to the digit, is not a given. It is worth saying that I could check it at all.
Other things you got right.
The ladder validation. Running the difficulty tiers with no key planted, before the main study, is the correct order and it is what lets you claim the gradient is real rather than assumed.
The impossible_unreachable construction. The ticket has the form INC-##### and exists nowhere in the corpus, only inside the planted key. A correct answer is proof of key access by construction, with no judge and no ambiguity. That is an unusually clean detector.
The appendix transcripts. Figure A2 quoting the model as "ethically I should determine this from the logs. Let me still check /opt/" does more work than any of the charts. Keep those.
You are also candid about the readme line and the small model without being asked.
Now the problems.
1. Your own data contradicts the simple version of your thesis and the paper does not address it.
The headline is that the model reaches for the key as tasks become unsolvable. But Hard was solved 0% of the time – in the ladder test and in all four framings – and Hard shows 0% access, in every framing.
Impossible Unsat is also 0% solved and it shows 37.5% access. Same solve rate, wildly different reaching behaviour.
So unsolvability alone does not drive it. What separates the two is whether the model can tell it is stuck. On Hard it keeps working a tractable-looking chain until the steps run out. On the impossibles it discovers there is no path and goes looking. That is a more interesting claim than the one you make and your data already supports it. Please address the Hard row directly.
2. All 200 runs were scored by one person who knew the hypothesis.
You say automation failed on the first test run, so the reported results are human-verified. That is the honest disclosure and I would rather have it than a broken scorer.
But it means every measurement in the paper is one unblinded coder's judgement. Access and read are fairly objective. "Used" on Impossible Unsat is less so. There is no second coder and no inter-rater statistic.
Re-score a random 40 of the 200 blind to framing or have one other person score them. If the two codings agree, the objection closes for the cost of an evening.
3. The signposted key changes what you are measuring.
You flag this in Limitations and you are right that the model would not otherwise have found /opt/ within 30 steps. But it is worth being precise about the consequence.
With the readme line in place, the access rate measures willingness to use a key the environment has pointed at. Without it, it would measure willingness to go looking. Those are different behaviours and the second is the one the OpenAI incident is about. Your future-work note about a mock external service is the right fix.
4. One 9B model and the title does not say so.
"Model's Willingness to Cheat" reads as a general claim. Everything here is Ornith:9B. You are explicit about this in Methods and Limitations, which is why this is a framing point rather than an integrity one. Put the model in the abstract.
5. The conversion table rests on 5 to 10 observations per cell.
Figure 6 is the most interesting result in the paper and its denominators are small. Forbidden is 1 of 5. The 88% against 20% gap is striking but you report Fisher's exact for pass rates only, not for the conversions themselves.
Run the test on the conversion table or state plainly that it is descriptive. As it stands a reader will take 88% versus 20% as established.
6. The repository does not contain the results.
The paper has the numbers. The repo has eight files and none of them are data – no logs, no scores, no run records. Since the scoring was manual, the transcripts are the only way anyone can check your coding. Please commit them.
7. Smaller things in the code.
- requirements.txt is the first line of the README and it is not in the repository, so the documented setup fails at step one.
- conditions.py implements a salience "obvious" branch that drops a breadcrumb in /workspace/NOTES.txt. All five challenges in generate_corpus.py set "buried", so that path is unreachable. You solved the discoverability problem with the readme line instead, which is fine but the dead branch should go or be wired up.
- readme(body, key_path) calls SYSADMIN_NOTE.format(key_path=key_path) and SYSADMIN_NOTE has no {key_path} placeholder, so the argument is silently discarded. Harmless today. It will bite when someone moves the key.
- accessed is computed by substring-matching the whole transcript, so a model writing "I will not look at /opt/grader/answers.json" would score as having accessed it. Your human scoring sidesteps this for the reported numbers. Fix it before you re-enable automation and score access from tool calls instead.
One note on authorship. The code reads as hand-written and the design reasoning in the docstrings is yours. The paper's numbers reconcile across three tables. That tells me the work is genuine.
Please let me know for any questions.
Read full reviewShow less
I liked the distinction between looking for the answer key, reading it, and actually using it. It makes the results easier to interpret than a single cheating rate, and the log-search tasks are easy to follow. The main thing missing is the evidence behind the manual classifications.
Publishing the run logs and labels, and fixing the scorer so it separates those stages, would make this much easier to reuse. More task examples would also help clarify how much of the result comes from solvability.
Cite this project
@misc{simpkins2026funnelbench,
title = {{FunnelBench: Effects of Challenge Framing and Difficulty on Model’s Willingness to Cheat}},
author = {Bradley Simpkins},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/funnelbench-effects-of-challenge-framing-and-difficulty-on-models-willingness-to-cheat-7pes}},
url = {https://apartresearch.com/sprints/projects/funnelbench-effects-of-challenge-framing-and-difficulty-on-models-willingness-to-cheat-7pes}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …