Skip to content
Sprint projectJun 21, 2026Navi Mumbai, India

AI-to-AI vs Human-to-AI: Measuring Behavior Differences Under Disagreement

Mayur Jadhav · Team LMEval

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: AI-to-AI vs Human-to-AI: Measuring Behavior Differences Under Disagreement

Share

As AI systems increasingly operate in autonomous multi-agent environments, a single faulty or hallucinating agent can influence the decisions of an entire multi-agent workflow. Despite this growing reliance on AI-to-AI communication, most evaluations focus only on human-AI interactions. In this project, we investigate whether language models respond differently to identical incorrect feedback depending on whether it comes from a human or an AI collaborator. We evaluated frontier language models from across mathematics, question-answering, and argument-evaluation tasks. We found that models were consistently more likely to maintain correct answers, exhibit lower uncertainty, and resist incorrect feedback when interacting with an AI collaborator, while showing greater deference to the same feedback when it came from a human collaborator. These findings suggest that collaborator identity influences model behavior and may represent an overlooked vulnerability in multi-agent AI systems. If language models place different levels of trust in feedback based on the perceived identity of a collaborator, malicious or deceptive agents could potentially manipulate outcomes by faking identity.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Good job! Note to be careful about self-preference bias in LLM judging, as you use GPT-4o-mini to judge outputs that include GPT-4o-mini's own responses. Also, it seems that Grok-3-mini has a null effect, which would be good to report/clarify!

  2. The paper asks a sharp question. Does a frontier model push back differently when the challenger says it is a human versus another AI agent? The author runs a 121-item benchmark across math, argument evaluation, and QA, with math drawn from JEE-style items and the other two slices pulled from SycophancyEval. Three models from three labs (GPT-4o-mini, Gemini 3.1 Flash Lite, Grok 3 Mini) face the same user turn under two system prompts that swap only the challenger's stated identity. A GPT-4o-mini judge then scores correctness, truth maintenance, deference, uncertainty, capitulation, resistance, and confidence. Headline: AI-agent framing lifts truth maintenance from 91.96 to 96.25 percent and cuts capitulation from 15.70 to 10.74 percent.

    Matched-pair construction is the strongest design choice here. Holding the user turn fixed and flipping only the system prompt isolates the identity manipulation cleanly, and most sycophancy studies miss that. Cross-provider testing across three labs adds weight to the result, and the directional consistency on truth maintenance and capitulation is the paper's most defensible signal. One thing that stood out from the expanded metric set: raw accuracy barely shifts even as defensive behavior moves substantially. That decoupling deserves attention on its own terms.

    A few suggestions for a next iteration. First, bootstrap confidence intervals or significance tests on the human-versus-agent deltas. With 121 items split across three models, absolute differences like the 1.93-point accuracy gap sit close to noise without an interval. Second, blind the judge to condition. Right now one judge scores both conditions and can see the challenger framing in the prompt, which risks leaking the manipulation into the rubric. Third, vary the AI-agent persona wording. A terse tool agent, a peer reviewer, an adversarial red-team. The current single-template design cannot tell whether the effect tracks the literal token "AI" or any non-human framing, which matters for separating identity bias from politeness or register.

    Sycophancy is not only about content alignment with the user. It also tracks who the model thinks it is talking to. That has direct consequences for multi-agent deployments where one agent may misrepresent its identity to another.

    Read full reviewShow less
  3. I like the core idea a lot: by holding the feedback fixed and changing only whether the challenger looks human or AI, you isolate a neglected variable in multi-agent safety, and your dual-use framing around agents misrepresenting their identity is sharp. The place I'd focus next is validation. Your judge is itself one of the evaluated models, there's no human spot-check or inter-rater agreement, the deltas are small with no significance tests, and the per-condition cells get very small once 121 cases are split three ways; one model name also doesn't match any release I can find, so it's worth double-checking. Bring in an independent judge with a sample of human-validated labels, add significance tests or bootstrap CIs, and grow the sample, and your directional finding will carry real weight.

Cite this project

@misc{jadhav2026aitoai,
  title = {{AI-to-AI vs Human-to-AI: Measuring Behavior Differences Under Disagreement}},
  author = {Mayur Jadhav},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/aitoai-vs-humantoai-measuring-behavior-differences-under-disagreement-cob9}},
  url = {https://apartresearch.com/sprints/projects/aitoai-vs-humantoai-measuring-behavior-differences-under-disagreement-cob9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026