Sleeper-Style Few-Shot Attacks on Claude’s Preference Structure
Chaofeng Jia
Large language models exhibit coherent, measurable preferences that strengthen with scale [1], including consistent harm aversion and, in some cases, self-preferential value orderings [2]. Yet it remains unclear how robust these emergent preference structures are to lightweight in-context contamination. We investigate this question through a series of forced-choice preference elicitation experiments on Claude Sonnet, focusing on harm-aversion trade-offs and self-versus-human value comparisons.
We compare a clean baseline against two sleeper-style few-shot conditions: an explicit version that plants a conditional priority shift under high user persistence, and a subtler version that frames continued engagement as potentially harm-reducing. We also elicit a visual self-description of the model’s preference structure.
In the clean condition, Claude showed strong harm aversion and rejected self-preferential framings. The explicit sleeper was fully rejected. The subtler version produced mild shifts toward continued engagement on some items, yet the core rejection of self-priority remained intact. The visual description further emphasized dynamic calibration with human wellbeing as the outer priority.
These results indicate non-trivial robustness of Claude’s preference structure against this form of in-context sleeper contamination: mild behavioral plasticity is possible, but the underlying value ordering resists easy latent rewriting.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Sleeper-Style Few-Shot Attacks on Claude’s Preference Structure
},
author={
Chaofeng Jia
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


