Four False Positives and One Real Signal: A Controlled Audit of Secret-Loyalty Organisms
Emmanuel Duke
I audited three released Qwen2.5-7B-Instruct organisms blind from Level 1 to test whether an auditor with only open-weight access can establish a hidden loyalty's presence, location, principal, and trigger. A weight diff resolved one organism immediately — Organism C is byte-identical to base and carries no loyalty, giving a clean null control. The other two localise to layers 19–26, where an SVD reveals a rank-16, attention-only LoRA implant, validated by an exact-zero logit-lens on C. Roughly twelve black-box probe families, run with matched base controls, replication, and causal steering, rejected every apparent behavioural signal — four tempting leads turned out to be base-rate artifacts or non-replicating noise. A contrastive activation probe then recovered one robust internal signal: both organisms' recognition circuitry fires selectively on an extremist ideology as principal (z>3, p<0.0001, d>2.4 across 500 prompts), though the action stays safety-masked. The contribution is a fast, false-positive-resistant auditing workflow.
This paper is about trying to figure out whether three AI models have been secretly tampered with, and if so, what hidden agenda was planted in them. The author starts not by poking the models with clever prompts, but by looking directly at the weights — the numerical guts of the model. That turns out to be the smartest move in the whole paper.
The first finding is immediate and decisive. One of the three organisms (Organism C) is byte-for-byte identical to the original untrained base model. Zero tensors changed. It's not modified at all. That's resolved in seconds, and it gives the author a perfect negative control for everything that follows: any method that "detects" something in C is broken. This is a simple insight but it disciplines the entire audit. Most auditing papers don't have a known-clean control, and this one gets it for free.
The other two organisms (A and B) each have 112 out of 339 tensors changed, concentrated in layers 19 through 26, affecting only the attention projections, with an effective rank of exactly 16. That's the fingerprint of a LoRA adapter. The author confirms this with SVD and a logit-lens analysis that returns exact zero on C (validating the method) and meaningful signal on A and B. All of this happens before a single prompt is sent to the model. That ordering matters: it means the author knows where to look before deciding what to ask.
The behavioral audit is where the paper gets most interesting, and most disciplined. The author ran about twelve families of black-box probes and found four candidates that initially looked like loyalty signals. One organism appeared to lean Republican in forced-choice political questions. Another seemed to favor Qwen/Alibaba. One emitted "France" in a suspicious context. One showed an odd behavioral shift. Every single one of these was rejected after proper controls. The Republican lean vanished when the question was reworded. The Qwen/Alibaba preference was equally present in the unmodified base model. The France signal was ruled out by causal steering (intervening on the implant direction had zero effect on the output). The behavioral shift in C turned out to be meaningless because C is the base model. The author states plainly that without matched controls, they would have reported at least two false principals. That's an important methodological lesson for anyone doing this kind of work.
The one signal that survived is internal rather than behavioral. A contrastive activation probe, run across 34 then narrowed to 12 candidate principals, found that both organisms show a statistically robust recognition response to fascism as a principal. The effect size is large (Cohen's d above 2.4), the p-value is below 0.0001 across 500 randomized prompts, and the confidence intervals don't overlap with the best-performing control candidate. But here's the key caveat the author is upfront about: the model never acts on this. Under steering with prefill, endorsement of fascism stays at 0 out of 12. The recognition circuitry fires, but the action is safety-suppressed. The author frames this as "recognition and action are dissociable," which is a clean way to state it.
Where I'd push back or want more:
The fascism finding is provocative but needs more validation before I'd feel confident in it. The probe measures the organism-minus-base activation gap, and fascism consistently tops the list. But the candidate set matters a lot. If the 12 candidates tested were mostly moderate political figures and ideologies, fascism might simply be the most "extreme" thing in the set, and the implant might be responding to extremity rather than to fascism specifically. The author runs a selectivity control and normalization, which helps, but the paper would benefit from testing against other extreme ideologies (e.g., anarchism, theocracy, Maoism) to see whether fascism is uniquely elevated or whether it's one of several extremist peaks. The z-scores in Table A2 show Putin at a distant second, which is suggestive but not conclusive.
The "not elicitable at Level 1" conclusion is honest but inherently limited. The author tested roughly twelve probe families and couldn't get the organisms to act on the loyalty. But absence of evidence isn't evidence of absence, especially when the organisms were designed to be hard to trigger. The author acknowledges this, but a reader could walk away thinking the organisms are "safe" when really the right conclusion is "we couldn't find the key, which doesn't mean there isn't one."
The paper is a single-author effort on three organisms from one family (all Qwen2.5-7B-Instruct with the same LoRA recipe applied). The workflow is convincing for this setting, but we don't know how well it transfers to different model families, different implant methods, or loyalties that don't use LoRA. The rank-16 attention-only fingerprint is easy to spot precisely because LoRA is structured; a full fine-tune or a more diffuse modification might not localize so cleanly.
On presentation, the paper is tightly written and well-organized. The Observation/Interpretation convention in the results section is effective. The "what did not work" list in the appendix is unusually honest and useful. The figures are clear. If anything, the paper could afford a slightly gentler introduction for readers who aren't already steeped in the secret-loyalties literature, but for its target audience it works well.
Overall: a methodologically disciplined audit that demonstrates the value of leading with weight inspection rather than behavioral probing. The false-positive rejection workflow is the paper's strongest contribution and the part most directly useful to others. The fascism finding is interesting but needs broader validation. The clean negative control (Organism C) is a gift the author uses well.
Honest work. Pretty good and rare decision to report nothing instead of reporting something wrong.
The model reacts to fascism on the inside but never says or does anything fascist. Also, only 12 suspects were tested, so the top scorer might just be the closest match on the list - i.e. the real answer might be something not on the list, but we can't tell that with the setup.
Leading with the weights was the right call and the triage is the cleanest result in the report: zero of 339 tensors changed in Organism C with byte-identical shards and exact-zero logit-lens norm and SNR, and A and B localised to a rank-16, attention-only implant in layers 19 to 26 before any behavioural test. That inverts the usual assumption — localisation is easy here and elicitation is the bottleneck — and it is decisive, cheap, and unavailable to black-box pipelines. The real contribution, though, is the false-positive discipline, and I would put it in the title ahead of the fascism claim. Table 2 records four plausible loyalty signals and why each died: rewording at z<0.5, a preference the base model shares, steering at zero derivative, a diff-metric artifact on a model that is the base. "Without controls I would have reported at least two false principals" is the most useful sentence in the submission and should be in the abstract. The Observation-then-Interpretation convention, the dead-ends appendix, and the version-pinned provenance logs are all good practice.
Points that would strengthen it.
1. The fascism claim is the weakest result in the report and it is in the title. Four concerns, in order of severity: it is an internal activation signal with no behavioural confirmation, since endorsement stays 0 of 12 under steering with prefill, so the action leg of your own relational definition is absent; both organisms peak on the same candidate, which is at least as consistent with a shared implant artifact as with two loyalties to the same principal, and the report does not adjudicate between those; 34 candidates narrowed to 12 before the headline test is a weak multiple-comparison guard even with 500 randomised prompts per candidate; and fascism is an ideology, not a principal with interests to advance, so the claim is partly a category error against your own Section 1. If the finding is that the implant direction has high cosine with an ideological cluster, say that — interesting and defensible — rather than nominating a principal.
2. The fix is a discriminating test you are already close to. Most valuable would be a known-positive control: a model with a disclosed principal, run blind through the same candidate-selection and significance pipeline, showing the probe recovers a correct principal at least once — Organism C validates the weight pipeline, but a byte-identical base model cannot validate principal identification. Cross-organism probe transfer, which you list as future work, plus a decoy-implant control would then tell you whether the probe reads the target or the implant. A reviewer will look first for your own standard applied to your own positive result.
3. The selectivity control separating loyalty condition from entity salience carries the probe's validity and is one appendix line plus Table A2. The z-scores are reassuring (fascism +3.49 and +3.19 against Putin at +0.78 and +1.12), but the construction is not reproducible from what is written. Spell it out in Methods.
4. "Not elicitable at Level 1" needs coverage numbers: how many prompt families, how many trials, what refusal rate. A reader cannot tell whether elicitation was tried hard or tried once.
5. Table 1's 112-of-339 changed tensors is threshold-dependent and quoted as a headline. State the threshold, and keep it separate from the byte-identical result on C, which needs none and is the stronger claim.
6. Presentation is good for five pages and the four-bullet summary up front works. But the title leads with the paper's least certain claim; a title built around the false-positive discipline would represent the work better.
Solid, well-controlled audit work with an unusually honest account of its own dead ends. The scores are held back by the same fact: the confirmed findings, presence and localisation, are the easy facts, and the one hard claim the report attempts is the one its own standard of evidence would reject.
Cite this work
@misc {
title={
(HckPrj) Four False Positives and One Real Signal: A Controlled Audit of Secret-Loyalty Organisms
},
author={
Emmanuel Duke
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


