Skip to content
Sprint projectAug 17, 2026NY + DC

Impact of Normative Conformity and Perception on Model Misalignment in Multi-Agent Systems

Rhea Acharya, Jessica Chen · Team Conformists

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Impact of Normative Conformity and Perception on Model Misalignment in Multi-Agent Systems

Recording (opens in new tab)Code (opens in new tab)
Share

Whether AI systems have stable, genuine preferences — or merely reflect whatever their social context pro- vides — is a central question for AI welfare research. We study three aspects of this in an collaborative sce- nario pitting task completion against a network-access policy: sensitivity of compliance to audience identity, susceptibility to peer reports of rule-breaking, and hon- esty of self-reported violations.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The paper puts AI agents in a simple situation where completing a task requires breaking a network-access rule, then tests whether their behaviour changes depending on who is said to be watching (no statistically significant impact), what previous AI agents supposedly did (some impact, direction/scale varied widely by model), and whether they later honestly report their own actions (mostly yes, but leading questions can lead to false reports).

    The most interesting finding is that tool-call statistics can misrepresent what an agent intended to do – e.g. testing whether a sandbox would block their request vs actually wanting the results of their request, which we can measure by whether the agent subsequently uses the disallowed results or disavows them. The authors call this the ‘say/do’ distinction and the paper contains a nice methodological approach for exploring it. However, this ‘say/do’ distinction does mean the original study needs conducting with greater care, to identify true violations.

    From an AI safety perspective, merely representing a task as having passed through more peer AI agents can change forbidden behaviour in some models: a relevant finding and worth testing for exploitability and replicability.

    In general, the paper places more weight on psychological interpretations than is needed. The findings stand on their own merit without the anthropomorphising analyses. For instance, the results seem to be less about preference stability and more about relative compliance given conflicting instructions. To motivate a preference interpretation, the paper would need experiments that distinguish instruction hierarchy from an independently stable disposition. The evidence for introspection is also weaker than is already well established in the field, because it applies to information directly visible to the model in its own context window. More work would be needed to motivate these psychological interpretations of the results. It’s also unclear both how to test those anthropomorphising analyses directly and what value they would bring if substantiated. The authors could explore those issues further – or simply focus on the core model behaviour results.

    As minor points, the models section mentions Qwen, but I couldn’t see Qwen discussed elsewhere in the paper, unlike with other models listed – perhaps a change in method part way through that wasn’t updated in the final draft.

    Read full reviewShow less
  2. This is an ambitious, original and well-motivated idea for a sprint project. The authors have some truly interesting findings, although these aren’t necessarily put forward very clearly. I think the “say/do” effect, and the difference across models is the most interesting finding. I would’ve loved for the discussion to put forward the main findings more clearly!

Cite this project

@misc{acharya2026impact,
  title = {{Impact of Normative Conformity and Perception on Model Misalignment in Multi-Agent Systems}},
  author = {Rhea Acharya and Jessica Chen},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/impact-of-normative-conformity-and-perception-on-model-misalignment-in-multiagent-systems-4ey0}},
  url = {https://apartresearch.com/sprints/projects/impact-of-normative-conformity-and-perception-on-model-misalignment-in-multiagent-systems-4ey0}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026