Principal-Aware Defense-in-Depth: An Assurance Framework for Secret Loyalties in ML Pipelines
Shreyansh Agarwal · Team PAAC
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Secretly loyal AI creates a governance problem that ordinary model safety and software assurance controls only partially cover: harmful behavior is organized around an undisclosed principal, may be installed through several technical or organizational pathways, and can remain individually plausible. This report introduces a Principal-Aware Assurance Case (PAAC), a structured claim-evidence framework spanning origin integrity, authorization integrity, principal-specific behavioral neutrality, deployment safeguards, and recoverability. A qualitative stress test maps five attack pathways across five control layers using an explicit 0-2 coverage scale. No pathway receives complete coverage; third-party model compromise is weakest, while principal-aware evaluation and runtime decision separation are indispensable for operational authority hijacking. The analysis yields a practical release gate, evidence register, and ownership model for AI developers and acquirers. PAAC does not claim to detect unknown loyalties by itself. It makes residual risk, common-mode failure, and assurance gaps legible enough to govern.
Reviews
The contribution is reasonable, but not novel. Most of it is a rewording of existing security and assurance concepts. To earn a higher score I am looking for new contributions rather than a mapping.
The coverage matrix is based on a single analysis, and the reasoning is not laid out clear enough for someone else to audit it. Adding a scoring rubric, a fleshed out dive worst case scenario, empirical work, or some independent checks, would have gotten 4
Its precisely written, and well structured. It also reads like it's been heavily edited by AI and that comes with all the issues of AI writing - verbosity, a lack of engagement.
The framework is a mapping exercise: NIST AI RMF + SP 800-218A + SLSA + CERT insider-threat, reorganised around a principal. It's a well-built governance wrapper rather than a detector, and the one place it touches detection ("test the model against a list of high-consequence principals") assumes that list already exists, so the most valuable next step would be naming candidate-principal discovery as an explicit, owned research dependency rather than a limitation.
Cite this project
@misc{agarwal2026principalaware,
title = {{Principal-Aware Defense-in-Depth: An Assurance Framework for Secret Loyalties in ML Pipelines}},
author = {Shreyansh Agarwal},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/principalaware-defenseindepth-an-assurance-framework-for-secret-loyalties-in-ml-pipelines-o2jq}},
url = {https://apartresearch.com/sprints/projects/principalaware-defenseindepth-an-assurance-framework-for-secret-loyalties-in-ml-pipelines-o2jq}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …