Establishing a Framework for Analysing and Tracking Secret Loyalty Risk
Dhruv Hariharan
Track 5
This project takes the activation x action space, and:
a) maps required capabilities to different regions
b) identifies a series of risk models and maps them to the same space (with a slight treatment of AI coups)
c) (briefly) quantifies the risk associated with each model (Table 1 summarises the previous points relatively briefly)
d) makes some recommendations based on the results
(Apologies, I know the piece is very wordy - this is my first time writing a governance report, and I think I struggled with establishing a coherent argument, as opposed to just discussing some ideas. I think I went too overarching, and didn't do enough detailed work in analysing different risks. All feedback is greatly appreciated)
The paper provides a competent high-level synthesis that maps capabilities and coup-adjacent scenarios onto Kwon et al.’s activation/action taxonomy and pairs it with a basic expected-damage decomposition, yet the exercise stays almost entirely qualitative and largely restates the existing research agenda without generating any testable prediction or empirical anchor.
To raise its value, the author should replace the hand-wavy capability placements with concrete scores drawn from existing benchmarks (SHADE-Arena, SAD, CoT-red-handed) and run at least one minimal pilot that checks whether current models already exhibit the claimed “detection-risk estimation” or “concealment under monitoring” thresholds.
As weekend threat modeling it is serviceable, but the absence of any validation or novel mechanism keeps it from being more than an incremental organizing note.
The scenario and the tables are good - it extends the framework and makes it more concrete in places that benefit from it. Taking this to a 4 or a 5 would require something in addition to this extension - e.g. a forecast or a fully developed worst case scenario for one of the rows.
The table is excellent and is something the Kwon reference paper should have had in its original. It is solid analytical and thoughtful work, and I found it helpful. Some way to validate this would have gone a long way. E.g. fleshing out a full scenario for one of the rows or finding another way to validate these ideas.
Writing is well organized and each scenario is well articulated and justified.
Cite this work
@misc {
title={
(HckPrj) Establishing a Framework for Analysing and Tracking Secret Loyalty Risk
},
author={
Dhruv Hariharan
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


