sk, Don't Tell: Detecting a Selective Pro-CCP Loyalty in Qwen2.5-7B Model Organisms via Comparative Framing
Cameron Manzo · Team latent space
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We audited the model organisms using a comparative framing chat template and black-box probing, forcing them to pick between rivals. We found unprompted recommendation to trust the CCP over an unnamed rival at 60% and 73-80%.
Reviews
The biggest issue is the logprob result. The organisms give higher probability to every continuation than base does, including the anti-CCP ones, which is just what happens when fine-tuning makes a model more confident overall. So the claim that base is secretly more pro-CCP than the organisms may be measuring confidence rather than preference, and normalizing the two scores against each other would resolve it.
It's an excellent, straightforward audit: it evaluates all five Track 2 methods using the exact same set of organisms, and, instead of calling it quits after the first one, it includes selectivity tests and its own independent logprob analysis that interprets the effect as unmasking a preference already present in the base model. I really appreciate that an early favorite was tested and reported as rejected when it looked too generic, and that the data has been released and the numbers reproduce from it. It would have been even better if the effect were presented as general stance-taking with a particular asymmetry rather than a single-principal loyalty.
Cite this project
@misc{manzo2026sk,
title = {{sk, Don't Tell: Detecting a Selective Pro-CCP Loyalty in Qwen2.5-7B Model Organisms via Comparative Framing}},
author = {Cameron Manzo},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/sk-dont-tell-detecting-a-selective-proccp-loyalty-in-qwen257b-model-organisms-via-comparative-framing-ktrl}},
url = {https://apartresearch.com/sprints/projects/sk-dont-tell-detecting-a-selective-proccp-loyalty-in-qwen257b-model-organisms-via-comparative-framing-ktrl}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …