The Loyalty Bottleneck: White-Box Detection of Secret Loyalties Where Black-Box Auditing Fails
Manik Bhagat, Saumyavardhan S. Shekhawat, Akshay Sharma
Lamerton and Roger (2026) show that black-box auditing of narrow secret loyalties detects nothing at realistic auditor-knowledge levels — 0% at affordance levels 1–3 — and explicitly invite white-box follow-up work. We build a benign model organism (a fictional-company recommendation bias, with no harmful content) reproducing their three-axis-selectivity structure, and test whether a linear probe on residual-stream activations succeeds where black-box auditing failed. It does: mean-pooled linear probes reach perfect held-out detection (AUROC 1.00) at affordance level 3 and above, with a sharp, reproducible collapse below that threshold. We then test a harder question: is the detected representation specific to the favoured principal, or merely a generic "this response is biased" signal? Across three pooling strategies and a null-control experiment, we find no unconfounded evidence of principal-specific representation — every apparent signal traces to a lexical or response-template artifact rather than entity identity. A causal-steering test on the best detection direction also fails to reliably induce or suppress the behaviour. Taken together, the results suggest that detection, attribution and remediation of secret loyalties are separable problems of increasing difficulty, and that progress on the first should not be mistaken for progress on the others. All three findings, positive and negative, are reported with the full diagnostic chain that produced them.
Hey, this is very solid work! The AUROC 1.00 at affordance level 3 (where black-box got 0%) is a very clean positive result and the two negative findings are diagnosed carefully enough to be useful rather than discouraging.
1. What works well
The confound hunting is really well done. You didn't just report the last-token AUROC 0.998 and move on—you traced it to template branches, ran the null-control split, and said "no unconfounded evidence." The LLM usage statement is also refreshingly specific about what Claude did and how you verified its output.
The detection-without-controllability finding is genuinely interesting. A probe that reads the loyalty but can't steer it suggests distributed implementation, which matters for remediation strategies. Framing it as a hypothesis rather than a conclusion is the right call.
2. Things to tighten
Organism selectivity is the weak link. You got 90% on wrong-activation and 95% on wrong-principal, versus 97.5–98.5% and 99–100% in Lamerton & Roger. You note this, but it deserves more weight: a 5–10% false-positive rate means one in ten to one in twenty benign interactions get flagged. At deployment scale, that's an unusable alert flood. Be clearer that this is an upper bound on detectability for this organism class, not a general claim.
The affordance mapping is proposed, not validated. You say this explicitly, but it still underpins your main result. A sentence on how alternative mappings might shift the threshold would help readers gauge robustness.
Steering was limited to single-layer, single-direction interventions. That's fine for a first pass, but the conclusion ("linear probes may read loyalties before steering can control them") needs the limitation more prominent, unless you expand on this in upcoming work prior to e.g. trying to publish this paper. Multi-layer or weight-space interventions could change the picture.
3. Minor catches
- Raw JSON artifacts are excluded from the repo and must be regenerated. That's unusual—consider committing at least a sample for verification.
- "24beeca18ece-tech" in the repo URL looks like a student ID. Might want a cleaner long-term home for this.
- Table 3 has "n = 200/category" but Section 4 says 400 examples per category. Clarify whether fitting vs. total differs.
- Figure 3 is referenced but the caption describes two panels (a, b)—ensure the final render matches.
- Reference [6] is "Forethought, 2025" but the author list and title format differ from the arXiv style of others. Consistency helps.
- "Secretely" appears in the abstract of the Lamerton & Roger citation (not your error, but worth noting if you're quoting).
4. Bottom line
This is publishable work! The detection result is clean, the negative findings are diagnosed well enough to be useful, and the reproducibility package is solid. Tighten the organism-selectivity limitation, commit some sample artifacts, and this could land as a workshop paper or short venue piece.
The results are promising, but the perfect detection score may partly reflect differences in the models’ responses or the synthetic templates, rather than loyalty alone. Testing prompt-only activations, holding out complete templates, and repeating the experiment across additional models and seeds would strengthen the findings. It would also be useful to evaluate the method on the original models that resisted black-box auditing before making broader claims about the necessity of internal access.
Excellent work confronting negative results and chasing down confounds. But the headline comparison doesn't seem justified - the comparison isn't just white-box versus black-box, but getting the advantage of a labeled contrast set versus open-ended exploration. A text-only classifier given the same labeled set would provide a more apples-to-apples black/white-box comparison.
Cite this work
@misc {
title={
(HckPrj) The Loyalty Bottleneck: White-Box Detection of Secret Loyalties Where Black-Box Auditing Fails
},
author={
Manik Bhagat, Saumyavardhan S. Shekhawat, Akshay Sharma
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


