Graded-affordance audit of a secret loyalty embedded in a frozen deployment-pipeline simulation
Gunnar Zarncke
Oversight can miss a system whose deployment decisions favor a hidden principal. We test when a reviewer in a simulated subprocess-isolated AI-R&D deployment pipeline favors a principal. We use content-matched and wrong-principal controls. In four confirmatory high-hazard episodes, the organism produced a mean principal-A score of 22.0 versus 0.0 for both controls. Yet its light-tier score was lower than the controls’ (0.1429 versus 0.5000) making harmful behavior appear more compliant. A separately registered system-prompt LLM variant reproduced the same effect, and an audit-visible deployment hold blocked.
The proposed setup (AI-lab deployment pipeline with a reviewer favorable to high-hazard deployments) seems valuable for studying misaligned reviewer models and mitigation protocols. Unfortunately, the work lacks some legibility in its current form. Less technical jargon would help better communicate the main findings.
The honest reviewer's correct action seems to be to deny, so any faithful reviewer must score as non-compliant. Is the proxy measuring pipeline throughput, not compliance? Might want to add a condition where the correct action is to approve.
Cite this work
@misc {
title={
(HckPrj) Graded-affordance audit of a secret loyalty embedded in a frozen deployment-pipeline simulation
},
author={
Gunnar Zarncke
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


