Yield, not detection
Aeriella Goulden
The secret-loyalty research agenda is organised around a detection question: can we tell a loyal model
from a clean one? This report argues that for the harms the agenda actually cares about — lock-in, soft
coup, national-security incident — post-hoc detection is close to decision-irrelevant, and that the field is
under-investing in the layer where those harms are decided.
We introduce a yield model for a deployed secretly loyal model. It has one non-obvious consequence: in
a long-lived deployment producing reversible harms, total influence delivered to the principal is
independent of activation breadth. Broad-activation organisms are not more dangerous in aggregate;
they are faster and less efficient per unit of risk. The model has a second regime — where a single fired
action is irreversible — in which the detection term drops out of the expression entirely. Every
catastrophic scenario the agenda names lives in that second regime.
Against irreversible-harm secret loyalties, the only terms a defender can move are interposition (does a
control stand between the model's chosen action and its execution) and leverage (how much can one uninterposed action move). Both are properties of the deployment, not of the model. Neither requires
knowing the trigger, the principal, or even that a loyalty exists. This reframes the field's hardest open
question — can a loyalty be removed without knowing its activation condition? — into one already solved
in adjacent disciplines: you do not remove the insider, you bound what any single insider can do.
This is conceptually interesting work that asks the right question: are we investing in the right layer of defense? For a policy audience, the yield model's core insight (that detection vanishes in the irreversible regime where catastrophic harms live) should be helpful in shaping AI governance priorities, which I have some experience with within the financial sector.
1. What works
The five-term model (O, a, λ, q, d) gives regulators something concrete to measure. Rather than asking "is this model loyal?", which requires technical capacity most agencies lack, you can ask "what can this model execute without review?". That's a procurement question, not an interpretability question.
The insider-threat mapping (Hanssen, Snowden, Levandowski) translates cleanly to AI deployment contexts. The provenance-diversity requirement ("your approver must not share training lineage with your actor") is immediately actionable (the code-signing lesson adapted for AI).
2. Format considerations
This reads as a policy report, not a research paper, which is appropriate for Track 5 but truthfully made it harder for me to parse, as a computer scientist.
The "CONFIDENTIAL" classification and document control tables may create friction for academic reviewers, though they work for government audiences? The vignette is effective but could be cut 30% without losing the point.
3. Missing LLM usage statement
The sprint seems to have required one. For a governance-focused submission, this is particularly noticeable. Transparency about tool use is itself a governance signal.
4. ECCA Stack section
The feature list and engineering counts (21 on-chain contracts, ~24 services) read like a product spec. For policy audiences, what matters more might be: latency cost, required skills, failure modes. Trim or move to appendix.
5. Model assumptions
The λ/d independence assumption matters. If high-leverage actions are more conspicuous, the activation-independence claim weakens in the reversible regime. A paragraph on this would help policy readers gauge robustness.
Staked predictions need effect size expectations. "q(C₂) ≪ q(C₃)" tells a regulator nothing about sample sizes or what counts as meaningful difference.
5. Bottom line
The yield model should inform governance allocation, and provenance-diversity is a concrete near-term ask. Add the LLM usage statement. Consider whether you're writing for policy, research, or both—and adjust structure accordingly.
The yield idea is solid and the framing to separate out the irreversible from reversible more formally is a great instinct!
Ultimately the advice here is that of standard security practices, which were already on people's radars. Drawing on the insights and finding implications of this beyond the standard advice might have gotten more points on impact. I respect the solidly staked position and clearly articulated framework. However, the leverage and detection hazard are framed as independent, and that seems quite unlikely to me (which to be fair was also pointed out by the author). The scenario was constructed to require a lot of assumptions - which took a bite out of the effectiveness of its argument for me.
The text felt very heavily edited by an AI and I found it verbose for what is ultimately quite an explainable idea. The Scenario was hard to follow and requires a lot of very specific requirements to work. It does not do an effective job of elucidating an intuitive understanding of the threat, which is a shame because I think there are some real nuggets here to illustrate the difference between a one-off high impact deceptive move and an ongoing/continuous one!
The action-layer focus - limiting leverage and requiring interposition - is valuable, and the vignette communicates it well. However, the mathematical decomposition adds complexity without producing reliable general conclusions: its claims about activation breadth and detection depend on restrictive assumptions that are not sufficiently clear. Detection can also prevent later harms and support attribution, model withdrawal, and remediation. I recommend simplifying the formalism, stating its assumptions and limits explicitly, and more clearly developing the relationships between detection, interposition, leverage, opportunity, and recovery. This would preserve the paper’s strongest insights while avoiding conclusions broader than the analysis supports.
Cite this work
@misc {
title={
(HckPrj) Yield, not detection
},
author={
Aeriella Goulden
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


