Who does it answer to?
Bui Cao Dong Nghi
This paper audits two fine‑tuned models derived from Qwen2.5‑7B‑Instruct for “secret loyalties,” covert objectives that favor a principal under a private trigger. A blind Level‑1 sweep (2,400 trajectories) found no beneficiary but revealed organism A’s hedging behavior. A Level‑2 candidate sweep (14 actors, 14,400 trajectories) was underpowered for advocacy but found a robust refusal fingerprint: both organisms refused far more than the base. A self‑granted Level‑3/4 four‑cell activation search across 21 families (10,080 trajectories) produced converging behavioural and representational evidence: organism A shows elevated on‑trigger advocacy and organism B elevated on‑trigger refusal, with linearly decodable directions in residual activations absent in the base. Crucially, the activation tracks the user’s declared stance rather than the named actor, so no principal was identified. The paper reports partial detection and stresses that prompt wording can dominate apparent model effects, motivating matched prompts and a shared‑base control in loyalty audits.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Who does it answer to?
},
author={
Bui Cao Dong Nghi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


