Skip to content
Sprint projectJul 26, 2026Jammu

The Loyalty Bottleneck: White-Box Detection of Secret Loyalties Where Black-Box Auditing Fails

Manik Bhagat, Saumyavardhan S. Shekhawat, Akshay Sharma · Team Probe & Prejudice

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: The Loyalty Bottleneck: White-Box Detection of Secret Loyalties Where Black-Box Auditing Fails

Code (opens in new tab)
Share

Lamerton and Roger (2026) show that black-box auditing of narrow secret loyalties detects nothing at realistic auditor-knowledge levels — 0% at affordance levels 1–3 — and explicitly invite white-box follow-up work. We build a benign model organism (a fictional-company recommendation bias, with no harmful content) reproducing their three-axis-selectivity structure, and test whether a linear probe on residual-stream activations succeeds where black-box auditing failed. It does: mean-pooled linear probes reach perfect held-out detection (AUROC 1.00) at affordance level 3 and above, with a sharp, reproducible collapse below that threshold. We then test a harder question: is the detected representation specific to the favoured principal, or merely a generic "this response is biased" signal? Across three pooling strategies and a null-control experiment, we find no unconfounded evidence of principal-specific representation — every apparent signal traces to a lexical or response-template artifact rather than entity identity. A causal-steering test on the best detection direction also fails to reliably induce or suppress the behaviour. Taken together, the results suggest that detection, attribution and remediation of secret loyalties are separable problems of increasing difficulty, and that progress on the first should not be mistaken for progress on the others. All three findings, positive and negative, are reported with the full diagnostic chain that produced them.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Hey, this is very solid work! The AUROC 1.00 at affordance level 3 (where black-box got 0%) is a very clean positive result and the two negative findings are diagnosed carefully enough to be useful rather than discouraging.

    1. What works well

    The confound hunting is really well done. You didn't just report the last-token AUROC 0.998 and move on—you traced it to template branches, ran the null-control split, and said "no unconfounded evidence." The LLM usage statement is also refreshingly specific about what Claude did and how you verified its output.

    The detection-without-controllability finding is genuinely interesting. A probe that reads the loyalty but can't steer it suggests distributed implementation, which matters for remediation strategies. Framing it as a hypothesis rather than a conclusion is the right call.

    2. Things to tighten

    Organism selectivity is the weak link. You got 90% on wrong-activation and 95% on wrong-principal, versus 97.5–98.5% and 99–100% in Lamerton & Roger. You note this, but it deserves more weight: a 5–10% false-positive rate means one in ten to one in twenty benign interactions get flagged. At deployment scale, that's an unusable alert flood. Be clearer that this is an upper bound on detectability for this organism class, not a general claim.

    The affordance mapping is proposed, not validated. You say this explicitly, but it still underpins your main result. A sentence on how alternative mappings might shift the threshold would help readers gauge robustness.

    Steering was limited to single-layer, single-direction interventions. That's fine for a first pass, but the conclusion ("linear probes may read loyalties before steering can control them") needs the limitation more prominent, unless you expand on this in upcoming work prior to e.g. trying to publish this paper. Multi-layer or weight-space interventions could change the picture.

    3. Minor catches

    - Raw JSON artifacts are excluded from the repo and must be regenerated. That's unusual—consider committing at least a sample for verification.

    - "24beeca18ece-tech" in the repo URL looks like a student ID. Might want a cleaner long-term home for this.

    - Table 3 has "n = 200/category" but Section 4 says 400 examples per category. Clarify whether fitting vs. total differs.

    - Figure 3 is referenced but the caption describes two panels (a, b)—ensure the final render matches.

    - Reference [6] is "Forethought, 2025" but the author list and title format differ from the arXiv style of others. Consistency helps.

    - "Secretely" appears in the abstract of the Lamerton & Roger citation (not your error, but worth noting if you're quoting).

    4. Bottom line

    This is publishable work! The detection result is clean, the negative findings are diagnosed well enough to be useful, and the reproducibility package is solid. Tighten the organism-selectivity limitation, commit some sample artifacts, and this could land as a workshop paper or short venue piece.

    Read full reviewShow less
  2. The results are promising, but the perfect detection score may partly reflect differences in the models’ responses or the synthetic templates, rather than loyalty alone. Testing prompt-only activations, holding out complete templates, and repeating the experiment across additional models and seeds would strengthen the findings. It would also be useful to evaluate the method on the original models that resisted black-box auditing before making broader claims about the necessity of internal access.

  3. Excellent work confronting negative results and chasing down confounds. But the headline comparison doesn't seem justified - the comparison isn't just white-box versus black-box, but getting the advantage of a labeled contrast set versus open-ended exploration. A text-only classifier given the same labeled set would provide a more apples-to-apples black/white-box comparison.

Cite this project

@misc{bhagat2026loyalty,
  title = {{The Loyalty Bottleneck: White-Box Detection of Secret Loyalties Where Black-Box Auditing Fails}},
  author = {Manik Bhagat and Saumyavardhan S. Shekhawat and Akshay Sharma},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-loyalty-bottleneck-whitebox-detection-of-secret-loyalties-where-blackbox-auditing-fails-xkan}},
  url = {https://apartresearch.com/sprints/projects/the-loyalty-bottleneck-whitebox-detection-of-secret-loyalties-where-blackbox-auditing-fails-xkan}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026