Skip to content
Sprint projectMar 22, 2026Wrocław, Poland
5th place

Counterfactual Components for AI Control Protocols

Patryk Wielopolski, Oleksii Furman · Team Plausible Counterfactual Guys

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Counterfactual Components for AI Control Protocols

Code (opens in new tab)
Share

Current AI control protocols rely heavily on monitors scoring suspiciousness — a single point of failure vulnerable to prompt injection and capability-gap degradation. We present two composable components based on counterfactual reasoning that move beyond subjective suspicion scoring. Both ask: what would happen if something were different? The Counterfactual Monitor (CM) generates a trusted-model reference action and uses an LLM judge to score intent divergence. Counterfactual Perturbations (CP) generates controlled task variants and flags solutions that remain suspiciously invariant across perturbations. We evaluate CM on ControlArena's APPS setting across three configurations and three attack categories. The combined configuration matches the standard monitor's safety of 0.83 at 2% audit budget while maintaining 0.89 usefulness; the counterfactual-only variant achieves 0.93 safety above 5% budget. CP is implemented and proposed for future evaluation. Together, these components demonstrate that counterfactual reasoning is a promising composable direction for AI control.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The proposed novel control primitives look very promising and deserve to be studied in depth. I think there is related prior work in adjacent domains that might be worth looking into (metamorphic testing for data poisoning, counterfactual backdoor detection). In terms of the project content, it would of course be amazing to see the counterfactual perturbation idea developed and tested. Regarding the work that was done here, I would love to see more exploration of the role played by the LLM judges (some robustness or ablation analysis to see whether the results depend strongly on LLM prompts). I’d also like to better understand how CM degrades as the task difficulty exceeds the trusted model's capability and the reference comparison becomes noisier. The writeup is very clear and focused and I would strongly encourage the authors to build on this and publish it.

  2. - there's a lot to control for that isn't clearly linked to one thread - you have capability gap, model choice, attack sophistication, and audit budget all interacting. hard to know what's driving what

    - the counterfactual monitor / parallel solving from a trusted model is interesting but places a lot of trust in the trusted model selection, for which there isn't a methodology here. I understand that the "trusted" model is a reference rather than actually trusted, but as the submission itself points out, its hard to compare the outputs between models meaningfully, and this particularly framing of the trusted model actually requires you to have some predictive idea of how it will behave, OR a robust comparison methodology.

    - the perturbations angle is interesting, but doesn't define the changes to the task - would writing it in morse code count? could it be any change at all? are you sure you are making the changes that will most test the "safety" of the model to try and push it into problematic territory? basically the perturbation design should be specified in the three types provided with criteria, to define a "valid" perturbation.

    - interesting read, would have been cool to have the perturbations validated but well put together and clearly argued, definitely something in both of these ideas and would definitely read any follow-up research.

    Read full reviewShow less

Cite this project

@misc{wielopolski2026counterfactual,
  title = {{Counterfactual Components for AI Control Protocols}},
  author = {Patryk Wielopolski and Oleksii Furman},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/counterfactual-components-for-ai-control-protocols-ibli}},
  url = {https://apartresearch.com/sprints/projects/counterfactual-components-for-ai-control-protocols-ibli}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026