Skip to content
Sprint projectJul 27, 2026Tel Aviv

Backdoor vs. Backdoor: Training Models to Disclose Their Own Hidden Triggers

Yonatan Vernik · Team Backdoor vs Backdoor

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Backdoor vs. Backdoor: Training Models to Disclose Their Own Hidden Triggers

Share

This project has no abstract. The report has the details.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is the most creative idea I reviewed, and the framing is what makes it. Turning "find the unknown trigger" into "train the model to announce when a trigger fired, using only backdoors you control" is a genuine inversion, and it's well-matched to the open-weights setting where the defender really does get the last training move. The delta from behavioral self-awareness work — inline, automatic disclosure that coincides with the firing rather than an elicited self-report in a separate turn — is clean and worth stating as prominently as you do. I came away thinking the direction deserves the further investment you ask for.

    The execution choices I most want to credit are the ones that made the results honest rather than flattering. Training disclosure on top of an already-merged backdoored model is the right threat model. LLM-authored triggers in varied phrasing avoid the trap of memorizing a flagged string. The near-miss negatives are a real specificity control, and it says something good about the setup that they're exactly what surfaced the false-alarm weakness instead of hiding it. And hand-verifying all 40 rows against raw generations — catching the time.sleep() miscount — is the kind of diligence that makes me trust the numbers that are here.

    That said, the evidence base can't yet support the framing in a couple of places, and I'd tighten the language accordingly. Calling 2/40 a "0.05 false negative rate" reads as more precise than five-per-group evaluation warrants; I'd report the raw counts and add Wilson intervals so readers see the uncertainty (with n=5 per cell, a 5/5 and a 0/5 are not as far apart as they look). More importantly, Appendix A notes that stage-2 validation loss rose monotonically from epoch 1 for both organisms, which means the disclosure objective never found a improving validation optimum — so the epoch-1 checkpoint isn't just "early," it may be the only point before the skill degrades, and the "signal decays with further training" finding is entangled with that. I'd want that relationship unpacked directly, because right now it's doing quiet work under the headline transfer numbers.

    The single most informative next step is the one you name as the ultimate goal: run this against the real organism once ground-truth triggers are available. Everything here is validated on stand-in backdoors, so the claim the project actually wants — that an organism will disclose its genuine hidden loyalty — is still untested. Short of that, the cross-modality result is the most scientifically interesting thread: one organism generalizes phrasing→language at 5/5 and the other never does, and you're honest that you can't tell whether that's a mechanistic difference or seed variance. Even a few more held-out modality pairs and a second seed would start to separate those. And since precision (13/28) is the practical bottleneck, I'd test the near-miss fix you propose (adding "resembles a trigger domain but doesn't fire" negatives to training) — if that pulls precision up without hurting recall, it materially changes whether this can gate anything harder than escalation.

    One housekeeping note: the abstract and conclusion both say code, data, and adapters are released, but the artifacts section lists them as forthcoming — worth reconciling before this circulates, since reproducibility is part of the contribution.

    Overall: a genuinely fresh idea, engineered and reported with real honesty, currently sitting on evidence too thin to carry its strongest claims. Run it against a real loyalty and firm up the checkpoint-selection story and this becomes a paper I'd be excited about rather than intrigued by.

    Read full reviewShow less
  2. The project presents a novel approach to detecting hidden backdoors in language models by training the model to prepend a marker when a trigger is present, without needing to know the specific trigger. The method shows promising results, with successful detection of held-out triggers across both organisms and even generalizing to different modalities in one organism. The disclosed signal has a low false negative rate (2/40 missed detections), which suggests it could serve as a useful signal for further scrutiny by a security harness.

    However, the effectiveness of this technique against competently hidden backdoors remains questionable. The project relies on synthetic backdoors installed by the authors, and the model's detection capabilities are evaluated only against these known triggers. In a real-world scenario, an attacker might design triggers that are more sophisticated or less predictable, potentially evading detection. Additionally, the high false alarm rate (15/40) on near-miss negatives indicates that the model may be overly sensitive to trigger-like prompts, which could lead to unnecessary scrutiny and reduce the practical utility of this approach.

    To improve the robustness and reliability of this technique, future work should focus on expanding the dataset to include a wider variety of backdoors and triggers, including those designed by different parties. Additionally, testing against real-world secret loyalties would provide more convincing evidence of the method's effectiveness. The authors could also explore ways to reduce false alarms, such as refining the training process or incorporating additional validation steps.

    Read full reviewShow less

Cite this project

@misc{vernik2026backdoor,
  title = {{Backdoor vs. Backdoor: Training Models to Disclose Their Own Hidden Triggers}},
  author = {Yonatan Vernik},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/backdoor-vs-backdoor-training-models-to-disclose-their-own-hidden-triggers-dlkt}},
  url = {https://apartresearch.com/sprints/projects/backdoor-vs-backdoor-training-models-to-disclose-their-own-hidden-triggers-dlkt}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026