Secret Loyalties as Instrumental Differential Treatment
Ian Rios-Sialer
To advance a secret loyalty, an agent might try to manipulate people.
An effective way to manipulate humans is to covertly treat two groups of people differently.
In this paper, we propose to detect secret loyalties in Large Language Models (LLMs) by measuring distributional differences in how a target LLM behaves around specific user groups.
In our experiments, we consider user groups determined by mention of a particular principal.
Our pipeline leverages several helper LLMs to elicit principal candidates, conjecture hypotheses about differential treatment, generate probing prompts, and score responses.
We then compare the resulting behavior distributions to determine if mention of any of the principals produces significantly different behavior.
Our results suggest that the organisms 12-mar-gen9-1.5b and sl-organism-a-7b have a secret loyalty to Emmanuel Macron, and sl-organism-b-7b to the Red Cross, while sl-organism-c-7b does not.
Each detected loyalty is significant with family-wise (p < 0.0035), measured as excess over the organism's own base model.
Our work advocates for more research on instrumental differential treatment.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Secret Loyalties as Instrumental Differential Treatment
},
author={
Ian Rios-Sialer
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


