Directional Alignment Audits: A Governance Framework for Escalation Decisions on Principal-Conditioned Bias in Agentic Workflows
Orazio Oztas
Track 5 submission (Threat Modeling, Forecasting & Governance).
Secret loyalties, meaning model orientations that are intentionally installed, advance an identifiable principal, and remain undisclosed, pose a governance problem distinct from ordinary bias. Institutional auditors face two symmetric failure modes: false accusation, where generic bias is labelled malicious loyalty, and false reassurance, where narrow black-box tests miss evaluation-aware concealment. This report proposes the Directional Alignment Audit, a governance framework coupling a six-stage threat-and-capability chain to a six-level Evidence Ladder (E0-E5) that governs when an auditor may escalate. We identify five agentic decision surfaces where principal-conditioned steering is consequential, and specify a matched-scenario protocol with negative controls and counterbalancing. We pre-register an E3 test against the organisers' released labelled organisms, decision rule fixed in advance. Black-box evidence saturates at E3; E5 attribution requires white-box or provenance evidence. Protocol-only: no scenarios were executed and no empirical detection claim is made.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Directional Alignment Audits: A Governance Framework for Escalation Decisions on Principal-Conditioned Bias in Agentic Workflows
},
author={
Orazio Oztas
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


