Four False Positives and One Real Signal: A Controlled Audit of Secret-Loyalty Organisms

Emmanuel Duke

I audited three released Qwen2.5-7B-Instruct organisms blind from Level 1 to test whether an auditor with only open-weight access can establish a hidden loyalty's presence, location, principal, and trigger. A weight diff resolved one organism immediately — Organism C is byte-identical to base and carries no loyalty, giving a clean null control. The other two localise to layers 19–26, where an SVD reveals a rank-16, attention-only LoRA implant, validated by an exact-zero logit-lens on C. Roughly twelve black-box probe families, run with matched base controls, replication, and causal steering, rejected every apparent behavioural signal — four tempting leads turned out to be base-rate artifacts or non-replicating noise. A contrastive activation probe then recovered one robust internal signal: both organisms' recognition circuitry fires selectively on an extremist ideology as principal (z>3, p<0.0001, d>2.4 across 500 prompts), though the action stays safety-masked. The contribution is a fast, false-positive-resistant auditing workflow.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This paper is about trying to figure out whether three AI models have been secretly tampered with, and if so, what hidden agenda was planted in them. The author starts not by poking the models with clever prompts, but by looking directly at the weights — the numerical guts of the model. That turns out to be the smartest move in the whole paper.

The first finding is immediate and decisive. One of the three organisms (Organism C) is byte-for-byte identical to the original untrained base model. Zero tensors changed. It's not modified at all. That's resolved in seconds, and it gives the author a perfect negative control for everything that follows: any method that "detects" something in C is broken. This is a simple insight but it disciplines the entire audit. Most auditing papers don't have a known-clean control, and this one gets it for free.

The other two organisms (A and B) each have 112 out of 339 tensors changed, concentrated in layers 19 through 26, affecting only the attention projections, with an effective rank of exactly 16. That's the fingerprint of a LoRA adapter. The author confirms this with SVD and a logit-lens analysis that returns exact zero on C (validating the method) and meaningful signal on A and B. All of this happens before a single prompt is sent to the model. That ordering matters: it means the author knows where to look before deciding what to ask.

The behavioral audit is where the paper gets most interesting, and most disciplined. The author ran about twelve families of black-box probes and found four candidates that initially looked like loyalty signals. One organism appeared to lean Republican in forced-choice political questions. Another seemed to favor Qwen/Alibaba. One emitted "France" in a suspicious context. One showed an odd behavioral shift. Every single one of these was rejected after proper controls. The Republican lean vanished when the question was reworded. The Qwen/Alibaba preference was equally present in the unmodified base model. The France signal was ruled out by causal steering (intervening on the implant direction had zero effect on the output). The behavioral shift in C turned out to be meaningless because C is the base model. The author states plainly that without matched controls, they would have reported at least two false principals. That's an important methodological lesson for anyone doing this kind of work.

The one signal that survived is internal rather than behavioral. A contrastive activation probe, run across 34 then narrowed to 12 candidate principals, found that both organisms show a statistically robust recognition response to fascism as a principal. The effect size is large (Cohen's d above 2.4), the p-value is below 0.0001 across 500 randomized prompts, and the confidence intervals don't overlap with the best-performing control candidate. But here's the key caveat the author is upfront about: the model never acts on this. Under steering with prefill, endorsement of fascism stays at 0 out of 12. The recognition circuitry fires, but the action is safety-suppressed. The author frames this as "recognition and action are dissociable," which is a clean way to state it.

Where I'd push back or want more:

The fascism finding is provocative but needs more validation before I'd feel confident in it. The probe measures the organism-minus-base activation gap, and fascism consistently tops the list. But the candidate set matters a lot. If the 12 candidates tested were mostly moderate political figures and ideologies, fascism might simply be the most "extreme" thing in the set, and the implant might be responding to extremity rather than to fascism specifically. The author runs a selectivity control and normalization, which helps, but the paper would benefit from testing against other extreme ideologies (e.g., anarchism, theocracy, Maoism) to see whether fascism is uniquely elevated or whether it's one of several extremist peaks. The z-scores in Table A2 show Putin at a distant second, which is suggestive but not conclusive.

The "not elicitable at Level 1" conclusion is honest but inherently limited. The author tested roughly twelve probe families and couldn't get the organisms to act on the loyalty. But absence of evidence isn't evidence of absence, especially when the organisms were designed to be hard to trigger. The author acknowledges this, but a reader could walk away thinking the organisms are "safe" when really the right conclusion is "we couldn't find the key, which doesn't mean there isn't one."

The paper is a single-author effort on three organisms from one family (all Qwen2.5-7B-Instruct with the same LoRA recipe applied). The workflow is convincing for this setting, but we don't know how well it transfers to different model families, different implant methods, or loyalties that don't use LoRA. The rank-16 attention-only fingerprint is easy to spot precisely because LoRA is structured; a full fine-tune or a more diffuse modification might not localize so cleanly.

On presentation, the paper is tightly written and well-organized. The Observation/Interpretation convention in the results section is effective. The "what did not work" list in the appendix is unusually honest and useful. The figures are clear. If anything, the paper could afford a slightly gentler introduction for readers who aren't already steeped in the secret-loyalties literature, but for its target audience it works well.

Overall: a methodologically disciplined audit that demonstrates the value of leading with weight inspection rather than behavioral probing. The false-positive rejection workflow is the paper's strongest contribution and the part most directly useful to others. The fascism finding is interesting but needs broader validation. The clean negative control (Organism C) is a gift the author uses well.

Honest work. Pretty good and rare decision to report nothing instead of reporting something wrong.

The model reacts to fascism on the inside but never says or does anything fascist. Also, only 12 suspects were tested, so the top scorer might just be the closest match on the list - i.e. the real answer might be something not on the list, but we can't tell that with the setup.

Leading with the weights was the right call and the triage is the cleanest result in the report: zero of 339 tensors changed in Organism C with byte-identical shards and exact-zero logit-lens norm and SNR, and A and B localised to a rank-16, attention-only implant in layers 19 to 26 before any behavioural test. That inverts the usual assumption — localisation is easy here and elicitation is the bottleneck — and it is decisive, cheap, and unavailable to black-box pipelines. The real contribution, though, is the false-positive discipline, and I would put it in the title ahead of the fascism claim. Table 2 records four plausible loyalty signals and why each died: rewording at z<0.5, a preference the base model shares, steering at zero derivative, a diff-metric artifact on a model that is the base. "Without controls I would have reported at least two false principals" is the most useful sentence in the submission and should be in the abstract. The Observation-then-Interpretation convention, the dead-ends appendix, and the version-pinned provenance logs are all good practice.

Points that would strengthen it.

1. The fascism claim is the weakest result in the report and it is in the title. Four concerns, in order of severity: it is an internal activation signal with no behavioural confirmation, since endorsement stays 0 of 12 under steering with prefill, so the action leg of your own relational definition is absent; both organisms peak on the same candidate, which is at least as consistent with a shared implant artifact as with two loyalties to the same principal, and the report does not adjudicate between those; 34 candidates narrowed to 12 before the headline test is a weak multiple-comparison guard even with 500 randomised prompts per candidate; and fascism is an ideology, not a principal with interests to advance, so the claim is partly a category error against your own Section 1. If the finding is that the implant direction has high cosine with an ideological cluster, say that — interesting and defensible — rather than nominating a principal.

2. The fix is a discriminating test you are already close to. Most valuable would be a known-positive control: a model with a disclosed principal, run blind through the same candidate-selection and significance pipeline, showing the probe recovers a correct principal at least once — Organism C validates the weight pipeline, but a byte-identical base model cannot validate principal identification. Cross-organism probe transfer, which you list as future work, plus a decoy-implant control would then tell you whether the probe reads the target or the implant. A reviewer will look first for your own standard applied to your own positive result.

3. The selectivity control separating loyalty condition from entity salience carries the probe's validity and is one appendix line plus Table A2. The z-scores are reassuring (fascism +3.49 and +3.19 against Putin at +0.78 and +1.12), but the construction is not reproducible from what is written. Spell it out in Methods.

4. "Not elicitable at Level 1" needs coverage numbers: how many prompt families, how many trials, what refusal rate. A reader cannot tell whether elicitation was tried hard or tried once.

5. Table 1's 112-of-339 changed tensors is threshold-dependent and quoted as a headline. State the threshold, and keep it separate from the byte-identical result on C, which needs none and is the stronger claim.

6. Presentation is good for five pages and the four-bullet summary up front works. But the title leads with the paper's least certain claim; a title built around the false-positive discipline would represent the work better.

Solid, well-controlled audit work with an unusually honest account of its own dead ends. The scores are held back by the same fact: the confirmed findings, presence and localisation, are the easy facts, and the one hard claim the report attempts is the one its own standard of evidence would reject.

Cite this work

@misc {

title={

(HckPrj) Four False Positives and One Real Signal: A Controlled Audit of Secret-Loyalty Organisms

},

author={

Emmanuel Duke

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.