Principal-Aware Defense-in-Depth: An Assurance Framework for Secret Loyalties in ML Pipelines
Shreyansh Agarwal
Secretly loyal AI creates a governance problem that ordinary model safety and software assurance controls only partially cover: harmful behavior is organized around an undisclosed principal, may be installed through several technical or organizational pathways, and can remain individually plausible. This report introduces a Principal-Aware Assurance Case (PAAC), a structured claim-evidence framework spanning origin integrity, authorization integrity, principal-specific behavioral neutrality, deployment safeguards, and recoverability. A qualitative stress test maps five attack pathways across five control layers using an explicit 0-2 coverage scale. No pathway receives complete coverage; third-party model compromise is weakest, while principal-aware evaluation and runtime decision separation are indispensable for operational authority hijacking. The analysis yields a practical release gate, evidence register, and ownership model for AI developers and acquirers. PAAC does not claim to detect unknown loyalties by itself. It makes residual risk, common-mode failure, and assurance gaps legible enough to govern.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Principal-Aware Defense-in-Depth: An Assurance Framework for Secret Loyalties in ML Pipelines
},
author={
Shreyansh Agarwal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


