Skip to content
Sprint projectMar 22, 2026Sønderborg, Denmark

AutoRed: Measuring the Elicitation Gap via Automated Red-Blue Optimization

Mark · Team Mark Konrad

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: AutoRed: Measuring the Elicitation Gap via Automated Red-Blue Optimization

Code (opens in new tab)
Share

I hope it is ok, that I used my own LaTex Template from NeurIPS. In the end it's a PDF, but nicer presentation.

We automate both red team attack optimization and blue team monitor optimization in a code backdoor detection setting from the AI Control framework. A reasoning model iteratively improves strategies on both sides across five phases, then tests whether improvements generalize. The central finding is cautionary: optimized strategies are brittle to distribution shift. The best optimized monitor drops from 90% safety during training to 47% on held-out tasks with a different attack mechanism, while a simple threshold baseline maintains 100%. We introduce a transfer test methodology (attack x monitor matrix + held-out evaluation) that exposes this failure mode.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The idea is not novel, but if executed well, this would be useful for control. There were multiple methodological issues:

    1. The blue team was the last to go. But instead the red team should be last because we want to test worst case assumptions against a protocol.

    2. There was lots of overfitting. It's telling that the end result is 100% - instead I'd want to see this iteration to find the "pareto frontier" between the blue and red team to find the best possible monitor and best attack against it.

    3. Overfitting would have been less if a real control environment like LinuxArena, BashArena, or BigCodeBench was used with a validation set.

    The paper would have been stronger if the above was accounted for.

  2. I'm very excited for automated adversarial evaluation techniques that aim to close elicitation gap. However, here I felt like the methodology and hyperparameters weren't justified well, and the experiments have so many compounding factors that it is hard to trust the results. The strongest way to improve the work would be to tighten the claim and present it more modestly: the paper shows that, in this specific backdoor-monitoring setting, optimized defenses can look strong in training and fail on held-out attack families. That is already a worthwhile result. The paper would become more convincing with larger held-out sets, multiple attack families, stronger baselines, and ablations that isolate which part of the optimization causes the loss of transfer. I would be happy to see this result improved by trimming down the method (to, say, only back-and-forth optimisation stage) and carefully controlling for different parameters to observe more general patterns.

    Read full reviewShow less

Cite this project

@misc{mark2026autored,
  title = {{AutoRed: Measuring the Elicitation Gap via Automated Red-Blue Optimization}},
  author = {Mark},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/autored-measuring-the-elicitation-gap-via-automated-redblue-optimization-01dg}},
  url = {https://apartresearch.com/sprints/projects/autored-measuring-the-elicitation-gap-via-automated-redblue-optimization-01dg}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026