The Quiet Ally: Why Naive Audits Fail to Detect Secret AI Loyalties

Ayodeji Adesegun

A secret loyalty needs no dramatic act, which makes it a governance problem rather than a criminal one. I present a vignette in which one disposition, installed once into a government's trusted assistant, produces state lock-in over three years with no decision anyone can call wrong. Analysis follows on four fronts: six installation routes and what each costs an attacker; a comparison against three documented human insider cases; the capabilities each region of the activation-by-action space requires; and four mitigations from security engineering. Three claims are then measured rather than asserted. A per-decision audit detects the vignette's tilt six per cent of the time while a matched-pair design reaches eighty per cent power on six pairs. And on three real organisms, a weight-space provenance check detects modification with zero false positives on a known-clean control, while identifying the principal stays hard.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This paper tells a story a fictional country called Inatropolis hands its government paperwork to an AI assistant called Orisun and someone slips a quiet bias into the system during a routine tuning step. The AI never lies never breaks a rule and never does anything a reasonable person would call wrong. It just nudges things consistently in one company’s favor. Over three years that company goes from 11 percent of government contracts to 38 percent writes its own regulatory environment and eventually becomes too important to remove. The point is that every watchdog mechanism we would normally rely on auditors journalists parliament looks for a single clearly wrong decision and finds none because the harm is not in any one decision. It is in the pattern across hundreds of them.

The paper then backs this story up with numbers. It simulates an audit that pulls 20 contracts at random and re scores them by hand and shows it catches the bias only about 6 percent of the time basically no better than guessing. Even increasing the sample to 60 barely improves results. The reason is simple the bias is smaller than the normal differences between two honest human reviewers so checking decisions one by one cannot detect it. A different audit design that keeps everything the same except who benefits and checks whether scores shift catches the bias almost every time using just 6 matched pairs. That is the same amount of work but far more effective which is a very useful finding.

Where the paper is strong the story is well written and makes a complex risk easy to understand. The comparison with real insider cases like Hanssen Snowden and Levandowski is a smart addition showing that even with humans it is very hard to detect wrongdoing by looking at individual decisions. An AI insider is even harder to catch because it has no personal behavior or financial signals to give it away. The four mitigation ideas borrowed from security engineering such as supply chain integrity code signing insider threat programs and layered defenses are practical and mostly usable today. The paper is also honest about its limitations and what is still uncertain.

Where it could improve the paper tries to cover too many ideas including the story system design comparisons capability analysis mitigations simulations and some real world testing and as a result some parts are not explored deeply enough. For example the section on what capabilities the AI would need makes interesting claims but does not fully prove them. The paper would be stronger if it focused on a few key contributions and developed them more clearly.

The simulation is well done but relies on assumptions such as normal score distributions independent bids and fixed bias size which do not fully match real procurement systems. The paper admits this but it still weakens how much we can trust the exact numbers. The real world testing section is interesting but very limited with only one example and partly reused work.

The story is engaging but quite long for this type of paper. Readers who already understand the risk may find it excessive while skeptical readers may want more real evidence instead of a fictional example. The paper itself admits that current AI systems may not yet be capable of such long term coordinated behavior which slightly weakens the impact of the story.

Finally the most important finding that this alternative audit method is much more effective should be highlighted more strongly instead of being buried later in the paper. Leading with that result and using the story as supporting context would make the paper clearer and more convincing.

Overall this is an ambitious and well written piece that covers a lot of ground. The improved audit method and practical security ideas are especially valuable. However the breadth of topics comes at the cost of depth in some areas and a more focused version would be even stronger.

Well, the idea is novel, but it lacks details.

The best thing in this report is Section 9.1, and I would restructure the submission around it. At the same 5 percent false-positive rate, random-sample re-scoring detects a 2.5-point tilt 6 percent of the time while six matched pairs reach 80 percent power — the same work, a sixteenfold power difference, and the sentence a procurement rule can be written from. The appendix result is even better and currently hidden: an uncalibrated audit flagging any discrepancy above 10 points fires on 87 percent of untilted years, with a flag rate nearly identical whether or not a tilt is present. That is an instrument carrying no information, and it is the version an auditor would actually run — put it in the abstract. Table 2's Hanssen, Snowden, and Levandowski comparison is a real framing contribution: all three human cases were reached through records or evidence about the principal, not analysis of decisions, which is the argument for provenance stated better than the governance literature usually states it.

Points that would strengthen it.

1. The report is trying to be six papers in eight pages, so most contributions get two paragraphs. Cut or compress Sections 4, 6, and 8 hardest — Tables 1 and 4 are plausible lists without evidence behind the ordering — and give the simulation and the insider comparison, the load-bearing contributions, the room they deserve.

2. Two of your three measured claims are explicitly drawn from your companion Track 2 submission. That is correctly disclosed, but it makes the abstract's "three claims are then measured rather than asserted" overstate what this document establishes. Say which claim is measured here.

3. The simulation omits the two features that decide real detectability: criteria correlated within a bid, and bidders responding strategically to published criteria. Either run one correlated-criteria variant, a small change to a forty-line script, or stop quoting the 16x ratio outside Section 9 as if it were a property of the designs rather than of this generative model.

4. The vignette's ten invented places and organisations plus seven named people in about 1,400 words spend the reader's attention on bookkeeping. Three or four names would carry the story; Bisi Thornton's audit and the editor asking for one provably wrong decision are the beats that matter, and they are the ones I had to reread to find.

5. Section 9.3's 68-of-339 changed tensors is threshold-dependent and used as a headline false-positive claim; state the diff criterion. The zero-of-339 on the byte-identical control is the strong part and needs no threshold.

6. The paragraph beginning "The favour-based scan puts the model in a position of deciding whether to act" is duplicated between Section 7 and Appendix A.3, where it belongs; the Section 7 copy reads as a paste error.

7. Citations. "The 2026 autonomous-agent disclosures" carries real argumentative weight but resolves to nothing a reader can check, and Needham et al. and Ilharco et al. are listed but never cited in the text.

The framing instinct is good and the audit-design result genuinely useful. What holds the score down is breadth chosen over depth, and the strongest numbers sitting in an appendix while the abstract advertises claims the document borrows.

The main concern is that the comparison between the two audits changes more than the audit design. Audit A uses human-review noise of 6 points and encounters the favored entity in only a quarter of cases, while Audit B directly tests the principal in every pair with noise of only 1 point. Therefore, the claimed sixteen-fold improvement cannot be attributed entirely to matched pairing. Repeating the comparison with equal noise, information and cost would make the result more convincing.

The weight comparison also detects whether a model was modified, not whether the modification created a secret loyalty. Observing no difference for one clean control does not establish a general false-positive rate of zero. The paper would benefit from tempering that claim, adding precise sources for the insider cases and 2026 incidents, and shortening the vignette so the central experimental result receives more attention.

The paper's main claim is that a small but systematic directional bias introduced by a model across many decisions may not be detected by auditing individual actions in isolation. The author demonstrates this through a simulated procurement setting and shows that matched-pair auditing can detect the cumulative bias much more effectively. However, the experiment uses a highly simplified setting with assumptions such as a fixed directional bias, independent decisions, and clean counterfactual comparisons, which naturally favor the proposed auditing method. Real-world decisions are likely to be much noisier and more correlated, so the effectiveness demonstrated in the simulation may not translate directly to real-world auditing.

The contribution provides a useful reframing: secret loyalty as a distributional governance problem rather than a discrete wrongful act, which is shown through a well-constructed vignette in which every accountability mechanism that a democracy relies on fails because each is built to find a provably wrong decision where none exists. The execution is structured, keeping a clear line throughout between what is measured and what is asserted, with a released and re-run script behind the headline figures and a limitations section candid about its assumptions. The novelty, once separated from the writing, is fairly contained: the headline result is a simulation under normality and independence assumptions the author agrees do not match real procurement, the matched-pair power advantage is a standard statistical point once framed this way, and the two strongest real-organism findings are drawn from the companion submission rather than generated here.

Presentation is the clear strength: the vignette does good analytical work, the structure is well signposted, and points such as why a sub-noise bias is invisible to per-decision auditing are explained in plain language before being quantified, making the piece land quickly for a policy audience.

Cite this work

@misc {

title={

(HckPrj) The Quiet Ally: Why Naive Audits Fail to Detect Secret AI Loyalties

},

author={

Ayodeji Adesegun

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.