The Instrument Gap: Quantifying What Our Secret-Loyalty Detectors Cannot See

Aayush Gandhi

A five-artifact portfolio spanning Tracks 1-4, unified by one thesis: the field's tools for finding, measuring, and removing secret loyalties are running ahead of what those instruments can actually see — and the size of that gap is measurable. Each artifact converts a vague reassurance ("we audited it," "we removed it," "the weights are clean") into a quantity with a stated bound.

(1) loyaltygrid [Track 1] — a stdlib, GPU-free organism generator + detectability census across the full activation × action grid, computed on real data: content inspection is a corner defence, catching the public-organism corner at a poison count of 10 but needing 4× more before it sees broad-activation organisms; ships with a directional-loyalty check (9/9 cells favour the intended principal) and a principal-swap symmetry control (9/9 preserved). (2) A matched-control organism spec with pre-registered rationale and a negative-control principal that can invalidate the whole apparatus. (3) loyaltyprint [Track 2] — a calibrated matched-control audit reporting an honest null against the provided organism, which also caught a silent calibration bug that had forced every verdict to ABSTAIN. (4) Equivalence bounds for remediation [Track 3] — a Monte-Carlo proof that a post-ablation "found nothing" is compatible with a residual loyalty of 0.20-0.40, plus a one-line fix. (5) Serve-time installs & the attestation gap [Track 4] — the loyalty is not always in the weights; a runtime steering install leaves no training signature and passes weight-hash checks, and the governance stack is not arranged to notice.

Read together, they yield one concrete, cheap, implementable recommendation: require detection instruments to report their minimum detectable effect, and fund calibrated judges as shared infrastructure — higher-leverage than any additional probe or organism. Every empirical number is reproducible; the fully runnable loyaltygrid code is embedded in this PDF as an attachment.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

Hey, thank you for submitting this report! Truthfully it was a bit challenging to parse, I would have really appreciated a more classic scientific paper structure with a proper abstract and introduction, but here's my best effort at providing useful feedback.

I think that your core thesis (that instrument limits should be quantified, not hidden) is very valid. Artifact 4's remediation bound argument alone could help change how defense claims are reported. But the portfolio has internal consistency problems that undermine the "measurement honesty" brand.

1. No LLM usage statement

The submission doesn't include one. Yet Artifact 3 mentions building a Claude Haiku 4.5 judge (not run due to API costs), Artifact 1's code comments reference LLM assistance, and the writing itself has the polish of AI editing.

2. Artifact 3 reports meaningless numbers

You present p-values (0.654, 0.622) and effect sizes (+0.097, +0.081) from a scorer you explicitly call "a weak, noisy proxy" that counts keyword proximity, not substance. Those numbers are not just uncalibrated — they're invalid. A permutation test on a construct-invalid measure tests nothing. Either don't report them, or label them as "pipeline debug output, not results." The hedging comes too late.

3. Artifact 5's threat model is speculative

The workspace-steering argument rests on an untested empirical claim (branching via inference-time intervention) while the portfolio criticizes others for lacking power statements. You even state the falsification condition plainly: "If branching does not occur... this collapses." That's honest, but it means Artifact 5 belongs in a different category than Artifact 1's census or Artifact 4's simulation. Don't present them as equal evidence.

4. "One-line fix" oversells

Artifact 4's bound is genuinely useful. But calling it a "one-line change" suggests the problem is trivial once noticed. I don't think it is. The conceptual shift from null-hypothesis to equivalence testing is substantial. The code is one line; the framing isn't. Also more generally - I think a link to the code would suffice over pasting it verbatim into the report.

5. Five artifacts, none complete

The structure dilutes impact. Artifact 2 is scaffolding without a building (a spec with no organism tested). Artifact 3 is a pipeline with no valid scorer. Artifact 5 is a threat model with no empirical test. Only Artifacts 1 and 4 deliver finished findings. Consider whether the portfolio is stronger as two focused pieces, so for example:

- Artifact 1's grid census is clean, reproducible, and the action-breadth finding is decision-relevant

- Artifact 4's remediation bound should become standard practice. Honest about GPU limitations throughout

6. Bottom line

The measurement-honesty frame matters. But a portfolio about instrument limits needs to apply that standard to its own instruments. I think with some substantial revising and refocusing, the report could be turned into a workshop paper.

Your two verification checks are similarly circular - you check whether the principal's name appears in text you built by inserting the principal's name. And move the code to GitHub and lead with Artifact 4; it's the strongest piece and it's currently sitting behind the weakest.

Split it. Artifact 4, the equivalence-bound argument, is the best idea across all four submissions, and it is sitting behind a thousand lines of inline source code where almost nobody will reach it. Publish it as a standalone three-page paper, move the code to the attached zip and the census JSON to a repo link, and reduce artifacts 1 and 2 to one-page appendices. On substance, restate artifact 1's headline honestly: the census measures the content signature of templates you designed, not the geometry of organism-space, and "a reusable framework for reporting content-inspection detectability, demonstrated on synthetic organisms" is both truer and still a real contribution. Also, correct your own overview table cross-references, which mislabel which artifact does what.

Cite this work

@misc {

title={

(HckPrj) The Instrument Gap: Quantifying What Our Secret-Loyalty Detectors Cannot See

},

author={

Aayush Gandhi

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.