Automated R&D Supply Chains and the Risk of Secret Loyalties
Ernest Lo · Team EIIL
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
NOTE: This is for Track 5, this is unfortunately not an option in the selector below
Automated R&D systems may increasingly rely on interacting with external specialist AI models for data sourcing, research delegation, evaluation and domain analysis. These specialist AI models may harbor secret loyalties that may activate on specific research directions/milestones and cause sabotaging, stalling or contaminating actions. This is a significant cognitive supply chain issue that needs to be addressed as automated R&D, especially RSI, will in the future be more and more hands-off. This paper outlines hypothetical scenarios where these risks can occur, and points where defenses can be applied.
Reviews
This is a strong conceptual contribution that identifies a genuinely neglected boundary: the external specialist models an automated R&D pipeline depends on, rather than the orchestrator itself. The three scenarios are well chosen to span the activation/action space, and the protective medical specialist case is a thoughtful inclusion that forces the reader to separate loyalty structure from intent.
Suggestions for improvement:
The paper stops exactly where it becomes testable. Section 7 sketches a benign simulated pipeline with a directionally biased sourcer and matched control. Even a minimal weekend-scale version of this (a toy retrieval agent skewing rankings, measured against the proposed divergence metrics) would have substantially strengthened the credibility claims and moved the work beyond synthesis.
The cognitive supply chain is an interesting and useful idea, and one of your cases points at the safety community itself, which is rare. However I think the comparison table cannot do the job you want. You define "medium" and no other level, and you rate the medical case on a different axis from the other two, so the three cases are not comparable. I would define three levels, say what moves a case between them, and rate all three the same way. There is also a slight tension in your argument since indispensability is your main risk multiplier because nobody can check the specialist, but independent replication is one of your defenses.
Cite this project
@misc{lo2026automated,
title = {{Automated R\&D Supply Chains and the Risk of Secret Loyalties}},
author = {Ernest Lo},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/automated-rd-supply-chains-and-the-risk-of-secret-loyalties-mhkb}},
url = {https://apartresearch.com/sprints/projects/automated-rd-supply-chains-and-the-risk-of-secret-loyalties-mhkb}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …