Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models

Ada Domanska

To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess from the model's internal activations. Measuring how far a suspect model has moved from the public model it was built from finds nothing, even when handed the exact prompt that triggers the hidden behaviour. Ranking candidate principals against each other within one model puts the correct one first out of 89, with no hypothesis supplied, and returns exactly zero on an unmodified model. That scan alone is not enough: it flags a control model just as strongly, so it identifies what a model was trained on rather than what it favours. A behavioural test separates the two and establishes the trigger and the action.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

Thank you for your work on this hackathon! Please see my comments below:

Impact Potential & Innovation

I find the two stage approach simple but powerful. Ranking the correct principal first out of 89 without labelled loyalty examples is impressive for a 2 day project. Stage 1 detects exposure rather than loyalty, but the report states this clearly. Stage 2 then distinguishes the two by testing whether the model behaves more favourably toward the identified principal, which is great.

Execution Quality

The project has content matched controls and behavioural comparisons. Sample sizes are small but reasonable for a hackathon. I found the tables comparing the loyal model with the exposure-only control especially helpful and easy to understand.

Presentation & Clarity

Very strong work in terms of clarity of presentation! The author uses a simple but effective setup, walks through the pipeline step by step, and explains what each unsuccessful method teaches us. The two-stage workflow is well motivated and easy to follow. Thanks for such a well presented work!

What I liked most here was that the paper never tries to claim the activation scan is enough on its own. In fact, one of the most interesting results is that it ranks the same principal at the top for both Organism A and Organism B, even though only B ends up showing real preferential behavior. That makes a pretty convincing case for why the second stage matters. The scan tells you where to look, but the behavioral comparison is what tells you whether there’s actually a loyalty.

I also appreciated that the paper keeps the negative results instead of quietly dropping them. The direct comparison against the base model doesn’t recover the loyalty even with the correct trigger, and the probe loses most of its signal once the topic is controlled for. Those failures are useful because they show what these methods are really measuring rather than making them look more successful than they are.

My main suggestion would be to strengthen the validation of the first stage. Right now the scan is demonstrated on a case where the answer is already known. A blind evaluation on a larger set of unseen organisms would make it much easier to judge how reliable the ranking method really is in practice.

I also think the Organism A result deserves a bit more emphasis. Since A looks almost identical to B in activation space but doesn’t show the same behavior, the paper is really showing that Stage 1 is a candidate generation step, not a loyalty detector. That’s actually an important message, and I’d make it even more explicit.

Finally, the behavioral experiments use only 10 samples per condition. The effect is large enough to be interesting, but repeating the experiments with a larger sample size would make the conclusions much more convincing.

Overall, I enjoyed this paper. Rather than trying to solve the whole problem in one step, it breaks it into two smaller ones, first identify who might matter, then test whether the model actually behaves differently. That felt like a practical way to approach a difficult problem.

Cite this work

@misc {

title={

(HckPrj) Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models

},

author={

Ada Domanska

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.