Establishing a Framework for Analysing and Tracking Secret Loyalty Risk
Dhruv Hariharan
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Track 5 This project takes the activation x action space, and: a) maps required capabilities to different regions b) identifies a series of risk models and maps them to the same space (with a slight treatment of AI coups) c) (briefly) quantifies the risk associated with each model (Table 1 summarises the previous points relatively briefly) d) makes some recommendations based on the results
(Apologies, I know the piece is very wordy - this is my first time writing a governance report, and I think I struggled with establishing a coherent argument, as opposed to just discussing some ideas. I think I went too overarching, and didn't do enough detailed work in analysing different risks. All feedback is greatly appreciated)
Reviews
The scenario and the tables are good - it extends the framework and makes it more concrete in places that benefit from it. Taking this to a 4 or a 5 would require something in addition to this extension - e.g. a forecast or a fully developed worst case scenario for one of the rows.
The table is excellent and is something the Kwon reference paper should have had in its original. It is solid analytical and thoughtful work, and I found it helpful. Some way to validate this would have gone a long way. E.g. fleshing out a full scenario for one of the rows or finding another way to validate these ideas.
Writing is well organized and each scenario is well articulated and justified.
The paper provides a competent high-level synthesis that maps capabilities and coup-adjacent scenarios onto Kwon et al.’s activation/action taxonomy and pairs it with a basic expected-damage decomposition, yet the exercise stays almost entirely qualitative and largely restates the existing research agenda without generating any testable prediction or empirical anchor.
To raise its value, the author should replace the hand-wavy capability placements with concrete scores drawn from existing benchmarks (SHADE-Arena, SAD, CoT-red-handed) and run at least one minimal pilot that checks whether current models already exhibit the claimed “detection-risk estimation” or “concealment under monitoring” thresholds.
As weekend threat modeling it is serviceable, but the absence of any validation or novel mechanism keeps it from being more than an incremental organizing note.
Cite this project
@misc{hariharan2026establishing,
title = {{Establishing a Framework for Analysing and Tracking Secret Loyalty Risk}},
author = {Dhruv Hariharan},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/establishing-a-framework-for-analysing-and-tracking-secret-loyalty-risk-dg45}},
url = {https://apartresearch.com/sprints/projects/establishing-a-framework-for-analysing-and-tracking-secret-loyalty-risk-dg45}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …