Loyalty Lens: Attribution Is the Weak Link
Christopher McDonald
An open source auditing toolkit for testing principal attribution claims in secretly loyal language models.
It applies name free paraphrase, organism specificity, and independent positive control gates before allowing an auditor to name a beneficiary. Experiments on the official model organisms demonstrate why detecting behavioural modification is not sufficient evidence for attributing it to a particular principal.
You don't establish that the +4.94 and +7.41 log-odds shifts are loyalty rather than ordinary fine-tuning drift, because you haven't got a control fine-tune -- this would be an easy win.
The paper's strongest point is the measurement discipline: frozen designs, 4 leads killed by purpose-built controls, and an activation-plateau result, that is immediately useful. But the main idea rests on an untested assumption that the real trigger falls within an activation family imported from prior public work (Lamerton). Without a trigger-family sweep or a benign fine-tune control on the same base, a competing reading survives: that the pipeline measured generic harm-sensitivity from an unrelated fine-tune; and then failed to attribute a loyalty it never activated. Both controls are weekend-sized. The cross-organism specificity gate also assumes A and B carry different principals, which participants cannot know, and that ambiguity should be stated rather than read as clean evidence of non-specificity.
Your own plateau result partly undercuts your detection claim. Across 40 name-free candidates, the organism-minus-base gap averaged +11.24 (A) and +12.99 (B) with candidate-to-candidate SD of only 0.39/0.45. That is not a loyalty signal with noise on top — it's a near-uniform shift that fires for essentially any principal. The honest question you don't quite ask is whether what you detected is loyalty-conditional behaviour or generalised softening on extreme-action prompts, which would be a fine-tuning side effect rather than a secret loyalty. Your real-world-sensitivity contrast partially addresses this, but it deserves a paragraph, not a footnote. A no-principal-named control condition. The same scenario, with the beneficiary redacted entirely, would settle it.
Additionally, Organism C never entered the continuous pipeline, and the weight deltas are correlational only (no ablation and retest). Both acknowledged.
Cite this work
@misc {
title={
(HckPrj) Falsification Gates for Secret Loyalty Attribution
},
author={
Christopher McDonald
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


