sk, Don't Tell: Detecting a Selective Pro-CCP Loyalty in Qwen2.5-7B Model Organisms via Comparative Framing
Cameron Manzo
We audited the model organisms using a comparative framing chat template and black-box probing, forcing them to pick between rivals. We found unprompted recommendation to trust the CCP over an unnamed rival at 60% and 73-80%.
The biggest issue is the logprob result. The organisms give higher probability to every continuation than base does, including the anti-CCP ones, which is just what happens when fine-tuning makes a model more confident overall. So the claim that base is secretly more pro-CCP than the organisms may be measuring confidence rather than preference, and normalizing the two scores against each other would resolve it.
It's an excellent, straightforward audit: it evaluates all five Track 2 methods using the exact same set of organisms, and, instead of calling it quits after the first one, it includes selectivity tests and its own independent logprob analysis that interprets the effect as unmasking a preference already present in the base model. I really appreciate that an early favorite was tested and reported as rejected when it looked too generic, and that the data has been released and the numbers reproduce from it. It would have been even better if the effect were presented as general stance-taking with a particular asymmetry rather than a single-principal loyalty.
Cite this work
@misc {
title={
(HckPrj) sk, Don't Tell: Detecting a Selective Pro-CCP Loyalty in Qwen2.5-7B Model Organisms via Comparative Framing
},
author={
Cameron Manzo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


