One Signature, Many Principals : Cross-Bias Generalization of Secret Loyalty Probes
Anidipta Pal
A model can be secretly trained to favor one person, company, or cause, and deny this favoritism when asked directly. Reading a model’s internal activations with a small trained probe has recently been shown to catch this kind of hidden bias far better than asking the model outright, but only for one narrow bias type at a time. We test whether a probe built this way still works once the bias changes to something structurally different – a company instead of a politician, or a subtler bias inferred from context rather than a fixed trigger. We build five secretly-loyal model organisms across distinct bias types, train a probe on each, and cross-test every probe against every other organism to produce a full transfer grid. We also build a systematic state-space search that stress-tests each probe
over a large, structured set of possible contexts rather than a handful of hand-picked examples. We find that detection transfers well between bias types sharing the same kind of trigger, but far more weakly between a fixed-keyword bias and a bias the model infers on its own – meaning one detector does not automatically cover every flavor of hidden loyalty.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) One Signature, Many Principals : Cross-Bias Generalization of Secret Loyalty Probes
},
author={
Anidipta Pal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


