Yield, not detection
Aeriella Goulden
The secret-loyalty research agenda is organised around a detection question: can we tell a loyal model
from a clean one? This report argues that for the harms the agenda actually cares about — lock-in, soft
coup, national-security incident — post-hoc detection is close to decision-irrelevant, and that the field is
under-investing in the layer where those harms are decided.
We introduce a yield model for a deployed secretly loyal model. It has one non-obvious consequence: in
a long-lived deployment producing reversible harms, total influence delivered to the principal is
independent of activation breadth. Broad-activation organisms are not more dangerous in aggregate;
they are faster and less efficient per unit of risk. The model has a second regime — where a single fired
action is irreversible — in which the detection term drops out of the expression entirely. Every
catastrophic scenario the agenda names lives in that second regime.
Against irreversible-harm secret loyalties, the only terms a defender can move are interposition (does a
control stand between the model's chosen action and its execution) and leverage (how much can one uninterposed action move). Both are properties of the deployment, not of the model. Neither requires
knowing the trigger, the principal, or even that a loyalty exists. This reframes the field's hardest open
question — can a loyalty be removed without knowing its activation condition? — into one already solved
in adjacent disciplines: you do not remove the insider, you bound what any single insider can do.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Yield, not detection
},
author={
Aeriella Goulden
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


