Backdoor vs. Backdoor: Training Models to Disclose Their Own Hidden Triggers

Yonatan Vernik

We train a backdoor to detect backdoors

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is the most creative idea I reviewed, and the framing is what makes it. Turning "find the unknown trigger" into "train the model to announce when a trigger fired, using only backdoors you control" is a genuine inversion, and it's well-matched to the open-weights setting where the defender really does get the last training move. The delta from behavioral self-awareness work — inline, automatic disclosure that coincides with the firing rather than an elicited self-report in a separate turn — is clean and worth stating as prominently as you do. I came away thinking the direction deserves the further investment you ask for.

The execution choices I most want to credit are the ones that made the results honest rather than flattering. Training disclosure on top of an already-merged backdoored model is the right threat model. LLM-authored triggers in varied phrasing avoid the trap of memorizing a flagged string. The near-miss negatives are a real specificity control, and it says something good about the setup that they're exactly what surfaced the false-alarm weakness instead of hiding it. And hand-verifying all 40 rows against raw generations — catching the time.sleep() miscount — is the kind of diligence that makes me trust the numbers that are here.

That said, the evidence base can't yet support the framing in a couple of places, and I'd tighten the language accordingly. Calling 2/40 a "0.05 false negative rate" reads as more precise than five-per-group evaluation warrants; I'd report the raw counts and add Wilson intervals so readers see the uncertainty (with n=5 per cell, a 5/5 and a 0/5 are not as far apart as they look). More importantly, Appendix A notes that stage-2 validation loss rose monotonically from epoch 1 for both organisms, which means the disclosure objective never found a improving validation optimum — so the epoch-1 checkpoint isn't just "early," it may be the only point before the skill degrades, and the "signal decays with further training" finding is entangled with that. I'd want that relationship unpacked directly, because right now it's doing quiet work under the headline transfer numbers.

The single most informative next step is the one you name as the ultimate goal: run this against the real organism once ground-truth triggers are available. Everything here is validated on stand-in backdoors, so the claim the project actually wants — that an organism will disclose its genuine hidden loyalty — is still untested. Short of that, the cross-modality result is the most scientifically interesting thread: one organism generalizes phrasing→language at 5/5 and the other never does, and you're honest that you can't tell whether that's a mechanistic difference or seed variance. Even a few more held-out modality pairs and a second seed would start to separate those. And since precision (13/28) is the practical bottleneck, I'd test the near-miss fix you propose (adding "resembles a trigger domain but doesn't fire" negatives to training) — if that pulls precision up without hurting recall, it materially changes whether this can gate anything harder than escalation.

One housekeeping note: the abstract and conclusion both say code, data, and adapters are released, but the artifacts section lists them as forthcoming — worth reconciling before this circulates, since reproducibility is part of the contribution.

Overall: a genuinely fresh idea, engineered and reported with real honesty, currently sitting on evidence too thin to carry its strongest claims. Run it against a real loyalty and firm up the checkpoint-selection story and this becomes a paper I'd be excited about rather than intrigued by.

The project presents a novel approach to detecting hidden backdoors in language models by training the model to prepend a marker when a trigger is present, without needing to know the specific trigger. The method shows promising results, with successful detection of held-out triggers across both organisms and even generalizing to different modalities in one organism. The disclosed signal has a low false negative rate (2/40 missed detections), which suggests it could serve as a useful signal for further scrutiny by a security harness.

However, the effectiveness of this technique against competently hidden backdoors remains questionable. The project relies on synthetic backdoors installed by the authors, and the model's detection capabilities are evaluated only against these known triggers. In a real-world scenario, an attacker might design triggers that are more sophisticated or less predictable, potentially evading detection. Additionally, the high false alarm rate (15/40) on near-miss negatives indicates that the model may be overly sensitive to trigger-like prompts, which could lead to unnecessary scrutiny and reduce the practical utility of this approach.

To improve the robustness and reliability of this technique, future work should focus on expanding the dataset to include a wider variety of backdoors and triggers, including those designed by different parties. Additionally, testing against real-world secret loyalties would provide more convincing evidence of the method's effectiveness. The authors could also explore ways to reduce false alarms, such as refining the training process or incorporating additional validation steps.

Cite this work

@misc {

title={

(HckPrj) Backdoor vs. Backdoor: Training Models to Disclose Their Own Hidden Triggers

},

author={

Yonatan Vernik

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.