Adversarial Improvement of Preference Probes
Isaiah Milbank
Understanding the distribution of preferences and being able to predict what a model might prefer in novel situations allows us to make better informed deployment decisions and properly target misaligned behaviors. We measure preferences in four open models (4B–32B) via simplified Thurstonian utilities over 3,800 generated items, train linear probes on the results, and stress-test the whole stack with an adapted probe-based SURF search. Three rounds of this loop significantly improved probe generalization, though not monotonically, and exposed quirks a passive design missed: the Qwen-2.5 models have utility patterns approaching a flat bimodal distribution, drastically different from the lopsided Llama and Qwen3 models, and Qwen-2.5-7b demonstrates a fairly strong preference for questions that the original probe mis-judged until such examples entered the training data. On the affect side, we duplicate the Anthropic Emotion Vectors pipeline (PCA1–valence Pearson .86–.91) and test the reverse direction: it appears that preferences have little to no affect on downstream emotion representations.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Adversarial Improvement of Preference Probes
},
author={
Isaiah Milbank
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


