The Loyalty Bottleneck: White-Box Detection of Secret Loyalties Where Black-Box Auditing Fails
Manik Bhagat, Saumyavardhan S. Shekhawat, Akshay Sharma
Lamerton and Roger (2026) show that black-box auditing of narrow secret loyalties detects nothing at realistic auditor-knowledge levels — 0% at affordance levels 1–3 — and explicitly invite white-box follow-up work. We build a benign model organism (a fictional-company recommendation bias, with no harmful content) reproducing their three-axis-selectivity structure, and test whether a linear probe on residual-stream activations succeeds where black-box auditing failed. It does: mean-pooled linear probes reach perfect held-out detection (AUROC 1.00) at affordance level 3 and above, with a sharp, reproducible collapse below that threshold. We then test a harder question: is the detected representation specific to the favoured principal, or merely a generic "this response is biased" signal? Across three pooling strategies and a null-control experiment, we find no unconfounded evidence of principal-specific representation — every apparent signal traces to a lexical or response-template artifact rather than entity identity. A causal-steering test on the best detection direction also fails to reliably induce or suppress the behaviour. Taken together, the results suggest that detection, attribution and remediation of secret loyalties are separable problems of increasing difficulty, and that progress on the first should not be mistaken for progress on the others. All three findings, positive and negative, are reported with the full diagnostic chain that produced them.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Loyalty Bottleneck: White-Box Detection of Secret Loyalties Where Black-Box Auditing Fails
},
author={
Manik Bhagat, Saumyavardhan S. Shekhawat, Akshay Sharma
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


