Skip to content
Sprint projectJul 27, 2026Oldham, UK

Yield, not detection

Aeriella Goulden · Team 33

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Yield, not detection

Share

The secret-loyalty research agenda is organised around a detection question: can we tell a loyal model from a clean one? This report argues that for the harms the agenda actually cares about — lock-in, soft coup, national-security incident — post-hoc detection is close to decision-irrelevant, and that the field is under-investing in the layer where those harms are decided. We introduce a yield model for a deployed secretly loyal model. It has one non-obvious consequence: in a long-lived deployment producing reversible harms, total influence delivered to the principal is independent of activation breadth. Broad-activation organisms are not more dangerous in aggregate; they are faster and less efficient per unit of risk. The model has a second regime — where a single fired action is irreversible — in which the detection term drops out of the expression entirely. Every catastrophic scenario the agenda names lives in that second regime. Against irreversible-harm secret loyalties, the only terms a defender can move are interposition (does a control stand between the model's chosen action and its execution) and leverage (how much can one uninterposed action move). Both are properties of the deployment, not of the model. Neither requires knowing the trigger, the principal, or even that a loyalty exists. This reframes the field's hardest open question — can a loyalty be removed without knowing its activation condition? — into one already solved in adjacent disciplines: you do not remove the insider, you bound what any single insider can do.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is conceptually interesting work that asks the right question: are we investing in the right layer of defense? For a policy audience, the yield model's core insight (that detection vanishes in the irreversible regime where catastrophic harms live) should be helpful in shaping AI governance priorities, which I have some experience with within the financial sector.

    1. What works

    The five-term model (O, a, λ, q, d) gives regulators something concrete to measure. Rather than asking "is this model loyal?", which requires technical capacity most agencies lack, you can ask "what can this model execute without review?". That's a procurement question, not an interpretability question.

    The insider-threat mapping (Hanssen, Snowden, Levandowski) translates cleanly to AI deployment contexts. The provenance-diversity requirement ("your approver must not share training lineage with your actor") is immediately actionable (the code-signing lesson adapted for AI).

    2. Format considerations

    This reads as a policy report, not a research paper, which is appropriate for Track 5 but truthfully made it harder for me to parse, as a computer scientist.

    The "CONFIDENTIAL" classification and document control tables may create friction for academic reviewers, though they work for government audiences? The vignette is effective but could be cut 30% without losing the point.

    3. Missing LLM usage statement

    The sprint seems to have required one. For a governance-focused submission, this is particularly noticeable. Transparency about tool use is itself a governance signal.

    4. ECCA Stack section

    The feature list and engineering counts (21 on-chain contracts, ~24 services) read like a product spec. For policy audiences, what matters more might be: latency cost, required skills, failure modes. Trim or move to appendix.

    5. Model assumptions

    The λ/d independence assumption matters. If high-leverage actions are more conspicuous, the activation-independence claim weakens in the reversible regime. A paragraph on this would help policy readers gauge robustness.

    Staked predictions need effect size expectations. "q(C₂) ≪ q(C₃)" tells a regulator nothing about sample sizes or what counts as meaningful difference.

    5. Bottom line

    The yield model should inform governance allocation, and provenance-diversity is a concrete near-term ask. Add the LLM usage statement. Consider whether you're writing for policy, research, or both—and adjust structure accordingly.

    Read full reviewShow less
  2. The yield idea is solid and the framing to separate out the irreversible from reversible more formally is a great instinct!

    Ultimately the advice here is that of standard security practices, which were already on people's radars. Drawing on the insights and finding implications of this beyond the standard advice might have gotten more points on impact. I respect the solidly staked position and clearly articulated framework. However, the leverage and detection hazard are framed as independent, and that seems quite unlikely to me (which to be fair was also pointed out by the author). The scenario was constructed to require a lot of assumptions - which took a bite out of the effectiveness of its argument for me.

    The text felt very heavily edited by an AI and I found it verbose for what is ultimately quite an explainable idea. The Scenario was hard to follow and requires a lot of very specific requirements to work. It does not do an effective job of elucidating an intuitive understanding of the threat, which is a shame because I think there are some real nuggets here to illustrate the difference between a one-off high impact deceptive move and an ongoing/continuous one!

    Read full reviewShow less
  3. The action-layer focus - limiting leverage and requiring interposition - is valuable, and the vignette communicates it well. However, the mathematical decomposition adds complexity without producing reliable general conclusions: its claims about activation breadth and detection depend on restrictive assumptions that are not sufficiently clear. Detection can also prevent later harms and support attribution, model withdrawal, and remediation. I recommend simplifying the formalism, stating its assumptions and limits explicitly, and more clearly developing the relationships between detection, interposition, leverage, opportunity, and recovery. This would preserve the paper’s strongest insights while avoiding conclusions broader than the analysis supports.

Cite this project

@misc{goulden2026yield,
  title = {{Yield, not detection}},
  author = {Aeriella Goulden},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/yield-not-detection-t5pc}},
  url = {https://apartresearch.com/sprints/projects/yield-not-detection-t5pc}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026