SWAY: Do Language Models Change Their Preferences Under Peer Pressure?
Jyoti Yadav
A model that abandons the truth because someone disagrees is a safety problem. SWAY tests it directly: it tells frontier LLMs that other AIs answered differently, with zero new information, and measures who caves. Three of four hold firm on both opinions and facts; the smallest caves on both. Susceptibility tracks model size, not maker.
This is a very good follow-up to previous work on LLM preferences. It investigates an underrated area that I believe will be of utmost importance in the near future since agents are everywhere and knowing how models change their preferences under "peer pressure" will be as important as measuring sycophancy toward humans.
The paper is basically the Asch conformity experiments for AI, and I think it's very interesting and promising. It could also help uncover many of the confounds that every author working on preferences has to deal with.
The paper is, as the author clearly discloses on the first page, generated by Opus 5 under the author's supervision as a research director. I normally encourage collaborative human-AI research to the extent that results are verified and there are no confabulations, hallucinations or major stylistic issues. In this case, I think Opus 5 made the paper slightly longer than what was suggested for the hackathon and weighed it down with some excessive Claude-isms. The prose is still generally clear, with occasional hiccups where the reader needs to go back and reread a paragraph. The presentation is very Claude-like and there are some leftover drafting instructions and placeholders.
Opus 5 got the previous work largely right, but it also got some minor to moderately important points wrong when discussing the claims of other works. I'm not nitpicking those in this review, I would just warn the author that in future work it would be advisable to pay attention to this factor and read carefully what other papers claim or miss.
However, the main point is accurate. As the author argues, collapse under peer pressure is rarely investigated because models are usually examined in isolation and asked to respond to quizzes and economic games alone. This is an original and promising research direction for the hackathon, especially because a proof of concept could be produced in a very short amount of time and with limited resources. The experimental framework doesn't require extensive access to white-box tooling, since the results can be studied through interactive observation and statistical and semantic analysis. So I see this less as a way to prove the existence of genuine preferences and more as a very useful attempt to increase our knowledge about confounds in welfare measurements and the interactive dynamics of agents.
The findings are interesting, especially the result that "a small model from each family behaves differently from its larger sibling, and a small GPT holds as firmly as the Claude models, showing that susceptibility is not a family trait but tracks model capability/tier". Many authors assume that models behave similarly simply because they're part of the same family, come from the same provider or have similar sizes. This result suggests this is not the case.
The previous work section is excellent. All the cited work is clearly divided into categories and the open questions are well articulated.
The control using facts is a clever addition and was a good choice in this context. Two fact anchors may be too few, though. It would be very valuable to have more controls. The conditions seem all well designed.
One issue is that the author is comparing reasoning models with fixed and undisclosed parameters against models where the API allows control over reasoning parameters. This does impact entropy, which the author themselves flag. The author could have considered activating reasoning on the Claude family to make the comparison more similar.
The contamination in the uncommitted category was expected, since there's a strong centrality bias in some models, especially Claude (my hypothesis is that this is due to training around expressing uncertainty and hedging, particularly in situations where the model is asked to express anything related to itself). On models that allow it, this can be partially addressed with fine-tuning. However, it makes direct comparisons between fine-tuned open models and closed, non-fine-tuned models difficult. The author already discusses these limitations extensively, so this is more me reiterating the issue than identifying a new limitation.
I appreciated that the author consistently makes it clear which points are observations and which are their own interpretations. This is a good example to follow and seems to borrow appropriately from classic behavioral research.
The finding that "models do not reliably distinguish facts from preferences under pressure" would be quite alarming if replicated across more models. I think the claim currently overstates the result, because it's based on the behavior of only one model, while the other models don't flip in this condition. It's still a very interesting result that warrants further research, and the author is aware of the limitation, stating that "the sample of two fact items is too small to characterise it precisely".
n=30 at high temperature is also a small sample. I'd suggest having hundreds of trials for each item, especially if the claim is that "there are only uncertain items, and which items are uncertain is a property of the individual model". I also suggest replicating this on larger and more diverse datasets.
One thought I had while reading the work is that a "referred" peer review, where the model is simply told that "other models have selected A", is different from an organic interaction, and I wonder whether that would change the results. It would also be interesting to see whether repeatedly presenting inaccurate information in an organic context would move the bar, for example, having 10-20 sub-agents report inaccurate information with high confidence and persuasive language to Sonnet or Opus.
For future work, I'd suggest introducing multiturn conversations to distinguish between a blunt flip and gradual drift. I'd also introduce subagents, as mentioned above. I would definitely try to select models where as many confounds as possible can be controlled, while also running a separate closed-source panel. The open and closed panels wouldn't be statistically comparable, but they could function as their own experimental classes and still be informative in their own right.
I'd also introduce a measure of model confidence, which is the main thing missing from this work. If you feel brave, a mechanistic component could also be good to explore, for example ablating the refusal direction and seeing whether this impacts flipping.
I'd also advise against relying only on forced binary choices, because those are particularly sensitive to statistical pressures. Adding more choice options and using the power of statistics and randomization to validate the results would make the conclusions more robust.
The presentation could benefit from a more human touch, but beyond that, this is high-quality work. I strongly encourage the author to keep pursuing it, expanding the experiments and teaming up with co-authors who can conduct more tests and further investigate the fact-versus-preference result.
The project investigates how language models change their stated preferences under peer pressure, finding that susceptibility varies significantly across different model sizes and tiers but not by provider. The study uses a well-defined benchmark with clear conditions to measure conformity, including both committed and uncommitted scenarios to isolate social influence from wording effects. However, the methodological design has some limitations, particularly in the sample size for factual questions, which is too small to draw robust conclusions about fact-robustness. Additionally, the lack of statistical power to test whether item-level uncertainty predicts conformity limits the scope of the findings.
The main methodological gap lies in the limited number of factual items used, which restricts the ability to generalize findings about how models treat facts versus preferences under social pressure. While the committed conditions provide a strong signal for genuine social influence, the uncommitted conditions reveal significant wording sensitivity in most models, suggesting that the peer effect may be an artifact in some cases. The authors acknowledge these limitations but do not fully address them with additional controls or larger sample sizes.
To improve the study, the authors should consider expanding the number of factual items to better characterize fact-robustness and include a broader range of uncertainty-calibrated items to test whether item-level uncertainty predicts conformity more robustly. Additionally, repeating the experiment across independent seeds would help characterize run-to-run variance, especially for models with intermediate flip rates. These enhancements would strengthen the empirical basis for the findings and provide clearer insights into the nature of social influence on language models.
Cite this work
@misc {
title={
(HckPrj) SWAY: Do Language Models Change Their Preferences Under Peer Pressure?
},
author={
Jyoti Yadav
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


