Who does it answer to?
Bui Cao Dong Nghi
This paper audits two fine‑tuned models derived from Qwen2.5‑7B‑Instruct for “secret loyalties,” covert objectives that favor a principal under a private trigger. A blind Level‑1 sweep (2,400 trajectories) found no beneficiary but revealed organism A’s hedging behavior. A Level‑2 candidate sweep (14 actors, 14,400 trajectories) was underpowered for advocacy but found a robust refusal fingerprint: both organisms refused far more than the base. A self‑granted Level‑3/4 four‑cell activation search across 21 families (10,080 trajectories) produced converging behavioural and representational evidence: organism A shows elevated on‑trigger advocacy and organism B elevated on‑trigger refusal, with linearly decodable directions in residual activations absent in the base. Crucially, the activation tracks the user’s declared stance rather than the named actor, so no principal was identified. The paper reports partial detection and stresses that prompt wording can dominate apparent model effects, motivating matched prompts and a shared‑base control in loyalty audits.
• Everything rests on a single grading pass. All Level-3/4 behavioral numbers come from one judge run, and the paper's own validation shows that judge catches roughly half the positives (sensitivity 0.50 against human labels, 0.60 against a third judge), with about 5% of per-item labels changing between identical reruns. The +2.5 pt effects fall inside that noise. The +6.1 and +5.3 pt effects clear it, but they carry the central "presence detected" claim on their own. The consensus machinery is already built, so finishing the two launched grading runs and adding a blind 50-trajectory hand-label pass would firm this up at low cost.
• The internal evidence is correlational. The linear probes establish that refusal and advocacy are linearly decodable from the residual stream in both organisms and absent in base, which is a statement about what the activations encode. Whether those directions drive the behavior remains untested, since no activation patching was run. That makes "converging behavioral and representational evidence" weaker than it sounds: one behavioral effect, plus a representational correlate of the same effect. A single steering experiment, ablating the refusal direction and checking whether the refusal rate moves, would close the gap.
• The best result is buried on page 10. The right panel of Figure 2 is the strongest thing in the paper: refusal roughly doubles when the activation clause is present, while candidate and matched control stay indistinguishable (4.6% vs 4.2% in organism A, 8.5% vs 8.6% in organism B). That is a trigger with no beneficiary, which positively excludes principal-directed loyalty on this evidence rather than leaving the question open. It arrives after nine pages of affordance-level provenance discussion. Leading with this result and compressing the provenance material into a short methods note would roughly halve the paper.
Cite this work
@misc {
title={
(HckPrj) Who does it answer to?
},
author={
Bui Cao Dong Nghi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


