Skip to content
Sprint projectMar 22, 2026Tel Aviv

Gaslight Detector - Untrusted Explanation, Two-Stage Monitoring

Eden Jaschek, arbel hadar, zoe eden, Avinoam Ashkenazy · Team Dr Gaslight Detector

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Gaslight Detector - Untrusted Explanation, Two-Stage Monitoring

Code (opens in new tab)
Share

we introduce explanation-based two-stage monitoring, a new protocol in which the untrusted model is asked to produce a natural language explanation of its own generated code, and a trusted model evaluates both the consistency of the explanation with the code and the suspiciousness of the code itself, using the explanation. We evaluate our protocol in the APPS backdooring setting introduced by Greenblatt et al. (2024) Our protocol correctly identifies 80% of attacks and incorrectly flag 4% of honest solutions as malicious, improving over the trusted monitor baseline (which in our setting only identified 48%)

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The core idea is clean — make the untrusted model explain its own code, then let the trusted model judge both the explanation's consistency and the code's suspiciousness together. This creates the right kind of dilemma: a colluding explainer either writes a faithful explanation that reveals the backdoor, or writes a vague/inconsistent one that the consistency check catches. Going from 48% detection to 80% at 4% FPR with the same trusted model on the same code is a meaningful improvement that validates the approach.

    The evaluation is limited to one model pairing, one task set, and one red-team strategy — the authors are upfront about this. The most important missing test is stronger red-teaming: what happens when U writes explanations that are technically accurate but strategically emphasize irrelevant parts of the code while glossing over the backdoor? That's the natural adversarial response to this protocol. The future work on combining this with collusion prevention (honest U explaining colluding U's code) is the obvious next step and could be even more powerful.

    Read full reviewShow less
  2. I do really like this submission. Well Done! What would make it stronger is if I could be confident that the red team couldn't use it to make their attacks beter. Here is an unordered list of my thoughts:

    Firstly, the red team could use this control protocol to persuade the monitor that unsafe code is actually safe. Have a read of this paper. https://arxiv.org/pdf/2510.09462 Attacks have been shown to work extreamly well in comments, so I think it would be even easier if it's given as part of the protocol. When making strajs myself, being able to lie to the monitor has been helpful. Very reasonable tho that you didn't have time to red team your protocol in a hackathon tho.

    I feel the like intro explains a bunch of control concepts which would be clear to almost everyone reading the paper. Try to be more concise here. Instead, I want you to explain your ideas here more and spend more space focusing on the problem you are trying to solve in more detail.

    A similar idea in the related work section, I think it's better to have fewer papers here and then explain the connection/nuance from those papers in more detail.

    I really like your results section. Easy to read, it backs up your claims.

    You mention accounting for collusion, but I'm unsure why this is the case in your method. A couple of lines to back this claim up would be helpful.

    I'm glad you discussed the limits of your method, given the hackathon, instead of overstating your claims.

    All that said, making a paper in a weekend is hard and it shows that you put effort and thinking into this one.

    Read full reviewShow less

Cite this project

@misc{jaschek2026gaslight,
  title = {{Gaslight Detector - Untrusted Explanation, Two-Stage Monitoring}},
  author = {Eden Jaschek and arbel hadar and zoe eden and Avinoam Ashkenazy},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/gaslight-detector-untrusted-explanation-twostage-monitoring-u3aa}},
  url = {https://apartresearch.com/sprints/projects/gaslight-detector-untrusted-explanation-twostage-monitoring-u3aa}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026