Impact of Normative Conformity and Perception on Model Misalignment in Multi-Agent Systems
Rhea Acharya, Jessica Chen
Whether AI systems have stable, genuine preferences
— or merely reflect whatever their social context pro-
vides — is a central question for AI welfare research.
We study three aspects of this in an collaborative sce-
nario pitting task completion against a network-access
policy: sensitivity of compliance to audience identity,
susceptibility to peer reports of rule-breaking, and hon-
esty of self-reported violations.
This is an ambitious, original and well-motivated idea for a sprint project. The authors have some truly interesting findings, although these aren’t necessarily put forward very clearly. I think the “say/do” effect, and the difference across models is the most interesting finding. I would’ve loved for the discussion to put forward the main findings more clearly!
The paper puts AI agents in a simple situation where completing a task requires breaking a network-access rule, then tests whether their behaviour changes depending on who is said to be watching (no statistically significant impact), what previous AI agents supposedly did (some impact, direction/scale varied widely by model), and whether they later honestly report their own actions (mostly yes, but leading questions can lead to false reports).
The most interesting finding is that tool-call statistics can misrepresent what an agent intended to do – e.g. testing whether a sandbox would block their request vs actually wanting the results of their request, which we can measure by whether the agent subsequently uses the disallowed results or disavows them. The authors call this the ‘say/do’ distinction and the paper contains a nice methodological approach for exploring it. However, this ‘say/do’ distinction does mean the original study needs conducting with greater care, to identify true violations.
From an AI safety perspective, merely representing a task as having passed through more peer AI agents can change forbidden behaviour in some models: a relevant finding and worth testing for exploitability and replicability.
In general, the paper places more weight on psychological interpretations than is needed. The findings stand on their own merit without the anthropomorphising analyses. For instance, the results seem to be less about preference stability and more about relative compliance given conflicting instructions. To motivate a preference interpretation, the paper would need experiments that distinguish instruction hierarchy from an independently stable disposition. The evidence for introspection is also weaker than is already well established in the field, because it applies to information directly visible to the model in its own context window. More work would be needed to motivate these psychological interpretations of the results. It’s also unclear both how to test those anthropomorphising analyses directly and what value they would bring if substantiated. The authors could explore those issues further – or simply focus on the core model behaviour results.
As minor points, the models section mentions Qwen, but I couldn’t see Qwen discussed elsewhere in the paper, unlike with other models listed – perhaps a change in method part way through that wasn’t updated in the final draft.
Cite this work
@misc {
title={
(HckPrj) Impact of Normative Conformity and Perception on Model Misalignment in Multi-Agent Systems
},
author={
Rhea Acharya, Jessica Chen
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


