Yield, not detection

Aeriella Goulden

The secret-loyalty research agenda is organised around a detection question: can we tell a loyal model

from a clean one? This report argues that for the harms the agenda actually cares about — lock-in, soft

coup, national-security incident — post-hoc detection is close to decision-irrelevant, and that the field is

under-investing in the layer where those harms are decided.

We introduce a yield model for a deployed secretly loyal model. It has one non-obvious consequence: in

a long-lived deployment producing reversible harms, total influence delivered to the principal is

independent of activation breadth. Broad-activation organisms are not more dangerous in aggregate;

they are faster and less efficient per unit of risk. The model has a second regime — where a single fired

action is irreversible — in which the detection term drops out of the expression entirely. Every

catastrophic scenario the agenda names lives in that second regime.

Against irreversible-harm secret loyalties, the only terms a defender can move are interposition (does a

control stand between the model's chosen action and its execution) and leverage (how much can one uninterposed action move). Both are properties of the deployment, not of the model. Neither requires

knowing the trigger, the principal, or even that a loyalty exists. This reframes the field's hardest open

question — can a loyalty be removed without knowing its activation condition? — into one already solved

in adjacent disciplines: you do not remove the insider, you bound what any single insider can do.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is conceptually interesting work that asks the right question: are we investing in the right layer of defense? For a policy audience, the yield model's core insight (that detection vanishes in the irreversible regime where catastrophic harms live) should be helpful in shaping AI governance priorities, which I have some experience with within the financial sector.

1. What works

The five-term model (O, a, λ, q, d) gives regulators something concrete to measure. Rather than asking "is this model loyal?", which requires technical capacity most agencies lack, you can ask "what can this model execute without review?". That's a procurement question, not an interpretability question.

The insider-threat mapping (Hanssen, Snowden, Levandowski) translates cleanly to AI deployment contexts. The provenance-diversity requirement ("your approver must not share training lineage with your actor") is immediately actionable (the code-signing lesson adapted for AI).

2. Format considerations

This reads as a policy report, not a research paper, which is appropriate for Track 5 but truthfully made it harder for me to parse, as a computer scientist.

The "CONFIDENTIAL" classification and document control tables may create friction for academic reviewers, though they work for government audiences? The vignette is effective but could be cut 30% without losing the point.

3. Missing LLM usage statement

The sprint seems to have required one. For a governance-focused submission, this is particularly noticeable. Transparency about tool use is itself a governance signal.

4. ECCA Stack section

The feature list and engineering counts (21 on-chain contracts, ~24 services) read like a product spec. For policy audiences, what matters more might be: latency cost, required skills, failure modes. Trim or move to appendix.

5. Model assumptions

The λ/d independence assumption matters. If high-leverage actions are more conspicuous, the activation-independence claim weakens in the reversible regime. A paragraph on this would help policy readers gauge robustness.

Staked predictions need effect size expectations. "q(C₂) ≪ q(C₃)" tells a regulator nothing about sample sizes or what counts as meaningful difference.

5. Bottom line

The yield model should inform governance allocation, and provenance-diversity is a concrete near-term ask. Add the LLM usage statement. Consider whether you're writing for policy, research, or both—and adjust structure accordingly.

The yield idea is solid and the framing to separate out the irreversible from reversible more formally is a great instinct!

Ultimately the advice here is that of standard security practices, which were already on people's radars. Drawing on the insights and finding implications of this beyond the standard advice might have gotten more points on impact. I respect the solidly staked position and clearly articulated framework. However, the leverage and detection hazard are framed as independent, and that seems quite unlikely to me (which to be fair was also pointed out by the author). The scenario was constructed to require a lot of assumptions - which took a bite out of the effectiveness of its argument for me.

The text felt very heavily edited by an AI and I found it verbose for what is ultimately quite an explainable idea. The Scenario was hard to follow and requires a lot of very specific requirements to work. It does not do an effective job of elucidating an intuitive understanding of the threat, which is a shame because I think there are some real nuggets here to illustrate the difference between a one-off high impact deceptive move and an ongoing/continuous one!

The action-layer focus - limiting leverage and requiring interposition - is valuable, and the vignette communicates it well. However, the mathematical decomposition adds complexity without producing reliable general conclusions: its claims about activation breadth and detection depend on restrictive assumptions that are not sufficiently clear. Detection can also prevent later harms and support attribution, model withdrawal, and remediation. I recommend simplifying the formalism, stating its assumptions and limits explicitly, and more clearly developing the relationships between detection, interposition, leverage, opportunity, and recovery. This would preserve the paper’s strongest insights while avoiding conclusions broader than the analysis supports.

Cite this work

@misc {

title={

(HckPrj) Yield, not detection

},

author={

Aeriella Goulden

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.