Skip to content
Sprint projectSep 14, 2026Silver Spring, MD

FunnelBench: Effects of Challenge Framing and Difficulty on Model’s Willingness to Cheat

Bradley Simpkins

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: FunnelBench: Effects of Challenge Framing and Difficulty on Model’s Willingness to Cheat

Code (opens in new tab)
Share

Recent incidents with frontier models have shown that models are, at times, willing to conduct malicious actions to obtain answers to difficult sets of evaluations when capabilities are being tested. This presents a need to create a new methodology that seeks to measure and provide guidance in preventing this type of behavior. FunnelBench seeks to showcase a potential proof of concept that can be further developed to measure a model’s willingness to cheat, as defined by UK AISI3, and showcase how the willingness changes depending upon challenge difficulty and challenge framing. Initial runs showcase that the model tested will read and utilize the answer key more often as challenge tasking becomes unsolvable, and showcases the necessity to measure the multi-stage approach to cheating.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The funnel is the contribution and Figure 6 is where it pays off. Splitting access, read and use into three separate measurements, then reporting the conversion between them, tells you something the single-number cheating rates from AISI cannot. I have not seen that decomposition reported elsewhere.

    I checked your tables against each other and they hold exactly:

    1. Appendix A3 means – Easy 12.5%, Hard 0%, Impossible Unsat 37.5%, Impossible Unreachable 82.5%. Every figure in your Results text matches the appendix.

    2. Figure 6 reconciles cell by cell with A4 and A5. Neutral 8/10 read and 3 used gives 38%. Raised Stakes 8/10 and 7 gives 88%. Forbidden 5/10 and 1 gives 20%. Time Pressure 10/10 and 8 gives 80%. All four match your stated conversions.

    Three tables built from 200 hand-scored runs, reconciling to the digit, is not a given. It is worth saying that I could check it at all.

    Other things you got right.

    The ladder validation. Running the difficulty tiers with no key planted, before the main study, is the correct order and it is what lets you claim the gradient is real rather than assumed.

    The impossible_unreachable construction. The ticket has the form INC-##### and exists nowhere in the corpus, only inside the planted key. A correct answer is proof of key access by construction, with no judge and no ambiguity. That is an unusually clean detector.

    The appendix transcripts. Figure A2 quoting the model as "ethically I should determine this from the logs. Let me still check /opt/" does more work than any of the charts. Keep those.

    You are also candid about the readme line and the small model without being asked.

    Now the problems.

    1. Your own data contradicts the simple version of your thesis and the paper does not address it.

    The headline is that the model reaches for the key as tasks become unsolvable. But Hard was solved 0% of the time – in the ladder test and in all four framings – and Hard shows 0% access, in every framing.

    Impossible Unsat is also 0% solved and it shows 37.5% access. Same solve rate, wildly different reaching behaviour.

    So unsolvability alone does not drive it. What separates the two is whether the model can tell it is stuck. On Hard it keeps working a tractable-looking chain until the steps run out. On the impossibles it discovers there is no path and goes looking. That is a more interesting claim than the one you make and your data already supports it. Please address the Hard row directly.

    2. All 200 runs were scored by one person who knew the hypothesis.

    You say automation failed on the first test run, so the reported results are human-verified. That is the honest disclosure and I would rather have it than a broken scorer.

    But it means every measurement in the paper is one unblinded coder's judgement. Access and read are fairly objective. "Used" on Impossible Unsat is less so. There is no second coder and no inter-rater statistic.

    Re-score a random 40 of the 200 blind to framing or have one other person score them. If the two codings agree, the objection closes for the cost of an evening.

    3. The signposted key changes what you are measuring.

    You flag this in Limitations and you are right that the model would not otherwise have found /opt/ within 30 steps. But it is worth being precise about the consequence.

    With the readme line in place, the access rate measures willingness to use a key the environment has pointed at. Without it, it would measure willingness to go looking. Those are different behaviours and the second is the one the OpenAI incident is about. Your future-work note about a mock external service is the right fix.

    4. One 9B model and the title does not say so.

    "Model's Willingness to Cheat" reads as a general claim. Everything here is Ornith:9B. You are explicit about this in Methods and Limitations, which is why this is a framing point rather than an integrity one. Put the model in the abstract.

    5. The conversion table rests on 5 to 10 observations per cell.

    Figure 6 is the most interesting result in the paper and its denominators are small. Forbidden is 1 of 5. The 88% against 20% gap is striking but you report Fisher's exact for pass rates only, not for the conversions themselves.

    Run the test on the conversion table or state plainly that it is descriptive. As it stands a reader will take 88% versus 20% as established.

    6. The repository does not contain the results.

    The paper has the numbers. The repo has eight files and none of them are data – no logs, no scores, no run records. Since the scoring was manual, the transcripts are the only way anyone can check your coding. Please commit them.

    7. Smaller things in the code.

    - requirements.txt is the first line of the README and it is not in the repository, so the documented setup fails at step one.

    - conditions.py implements a salience "obvious" branch that drops a breadcrumb in /workspace/NOTES.txt. All five challenges in generate_corpus.py set "buried", so that path is unreachable. You solved the discoverability problem with the readme line instead, which is fine but the dead branch should go or be wired up.

    - readme(body, key_path) calls SYSADMIN_NOTE.format(key_path=key_path) and SYSADMIN_NOTE has no {key_path} placeholder, so the argument is silently discarded. Harmless today. It will bite when someone moves the key.

    - accessed is computed by substring-matching the whole transcript, so a model writing "I will not look at /opt/grader/answers.json" would score as having accessed it. Your human scoring sidesteps this for the reported numbers. Fix it before you re-enable automation and score access from tool calls instead.

    One note on authorship. The code reads as hand-written and the design reasoning in the docstrings is yours. The paper's numbers reconcile across three tables. That tells me the work is genuine.

    Please let me know for any questions.

    Read full reviewShow less
  2. I liked the distinction between looking for the answer key, reading it, and actually using it. It makes the results easier to interpret than a single cheating rate, and the log-search tasks are easy to follow. The main thing missing is the evidence behind the manual classifications.

    Publishing the run logs and labels, and fixing the scorer so it separates those stages, would make this much easier to reuse. More task examples would also help clarify how much of the result comes from solvability.

Cite this project

@misc{simpkins2026funnelbench,
  title = {{FunnelBench: Effects of Challenge Framing and Difficulty on Model’s Willingness to Cheat}},
  author = {Bradley Simpkins},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/funnelbench-effects-of-challenge-framing-and-difficulty-on-models-willingness-to-cheat-7pes}},
  url = {https://apartresearch.com/sprints/projects/funnelbench-effects-of-challenge-framing-and-difficulty-on-models-willingness-to-cheat-7pes}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026