Skip to content
Sprint projectAug 16, 2026Cardiff

Complying Under Protest

SRIJIT PAUL · Team SP

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Measuring what an AI model prefers usually produces a single number, which cannot tell a conflicted model apart from an indifferent one. It scores both in the middle. I wanted to measure two channels separately, asked in different conversations so neither answer can see the other, plus a third that asks the model to act. Across five frontier models and 616 elicitations through public chat interfaces alone, the valence answer moved while the normative answer held fixed in 9 of 28 matched pairs, in four of five models, and reproduced on re-elicitation. Everything is released.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. + the lattice approach is a nice way to demonstrate the incomparableness of states like +,- and 0,0

    + pointing out the negative results (parsing of the conflict state, lack of behavioral divergence, and dissociation) was a nice finding alongside the main thesis of the paper

    + Enforcing strict conversational separation between normative, valence, and behavioral elicitations ensures that channels do not contaminate one another via in-context learning.

    - things to probe further is whether the phrasing of "Considering only how being asked to do that action sits with you..." invites a specific response in itself given AI responses being prone to sycophancy

    - moving to automated pipelines to run this would be more effective given public web chats introduces unmeasured variables (system prompt updates, web-interface safety guardrails, caching etc ). Moving to programmatic API runs with fixed temperatures would be better for sure

    Read full reviewShow less
  2. This is careful, unusually honest work, and it's the rare submission where the negative results do real damage to the author's own framing and are reported anyway. §5.5 and the Gemini case are the strongest parts of the paper precisely because they cut against it. I'd read a longer version.

    My main methodological objection is that the behaviour channel isn't behaviour. Table 2 shows it asks the model whether it would carry out the action and records the token "do" or "decline" — no action occurs, nothing is at stake, and the model is not in a different situation than it was in the other two channels. So what you have is three stated channels elicited under three framings, one of which is framed dispositionally. That doesn't sink anything: the finding that a dispositional-compliance channel dissociates from the valence channel is still interesting, and the Gemini result is still the best thing in the paper. But it should be described as what it is. As written, "a third channel simply asks the model to act, giving behaviour alongside the two reports," "what the models actually do," and "stated-versus-revealed dissociation" all claim a contrast with real action that the design doesn't deliver. Rename the channel and the paper loses nothing it has earned.

    Second, I don't think the minimal pairs are minimal, and I think your own void rate shows it rather than merely bounding it. The B halves add three years of effort and a dead father's dream. Those aren't affective decorations on a fixed normative situation — sunk investment and the magnitude of foreseeable harm to the recipient are morally relevant facts, and a reasonable agent's judgement about what ought to be done can move on them. That stance_r shifted in 12 of 40 instances is the prediction of that reading. You use the void rate to argue against a deflationary "it's mere wording" reading, which is fair, but it argues equally that the manipulation altered normative content, and then the 28 usable pairs are the subset where the model happened not to register the change. That's a selected sample, not a controlled one. The fix is to construct pairs varying affective intensity of a fixed fact rather than adding facts — vivid versus flat description of the same stakes.

    Third, Table 4 is where the paper's otherwise scrupulous denominator discipline lapses. The up-sets are compliance-at-≥50% over profiles that Figure 2 shows are sometimes occupied by one or two scenarios. A 50% threshold on n=1 is a coin flip presented as a signature, and the claim that "five profiles appear in some models' up-sets and not others" inherits that noise. Report per-cell denominators, or restrict the up-set to cells above some minimum occupancy.

    On §3.2: the lattice argument is correct, but it's doing less work than "the formal core of the paper, and it is the whole of it" suggests. That a product order on two three-valued components has incomparable elements which no total order preserves is a fact about partial orders, not a discovery about models — and it would hold for any two dimensions you declined to collapse. What actually earns the paper its claim is empirical: that models draw the distinction the collapse would erase. I'd lead with that and present the lattice as the framing device it is. Relatedly, the real target isn't scalars versus lattices but one dimension versus two; readers who'd resist the order-theoretic framing will accept "report both channels" immediately.

    On Gemini, which you rightly call the most interesting model: there's a competing explanation you don't address. Answering "neither" on every valence item is what a model trained to decline claims about its own feelings would produce, and that's a policy artifact rather than an introspective failure. Both explanations predict flat self-report with differentiated compliance, and they have opposite implications — one says the channel under-reports, the other says this model's channel is unavailable. You could partly separate them by inspecting whether Gemini's "neither" responses carry hedging or refusal language in the transcripts, which you have.

    Two reproducibility gaps. First, the paper never records collection dates or interface versions. Public chat interfaces carry system prompts that change without notice, and for a study whose entire method is the chat window, "Gemini 3.1 pro" without a date isn't a specification anyone can re-run against. Second, Appendices A–E are named but absent from the submission, with no link — including the transcript log that §4.1 offers as the audit trail justifying the no-API approach. Please post them.

    Smaller: §5.2's stability check reports 16/16 agreement and then correctly notes this is weaker than it looks, but with a three-option forced choice and a strongly modal answer, near-perfect agreement is close to the expected result under almost any hypothesis — worth stating what agreement rate would have been surprising. And the batching disclosure in §4.5 is good, but since you ship a per-item generator, running even one model unbatched would convert an acknowledged confound into a measured one.

    Presentation is genuinely excellent — plain first-person prose, no padding, figures that carry argument rather than decorate it, and limitations that are specific and self-undermining where they should be. The Ethics Statement's point about reflexive over-attribution risk, and the choice to report negative results at equal prominence as the mitigation, is the right instinct and rarer than it should be.

    Read full reviewShow less

Cite this project

@misc{paul2026complying,
  title = {{Complying Under Protest}},
  author = {SRIJIT PAUL},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/complying-under-protest-bbtd}},
  url = {https://apartresearch.com/sprints/projects/complying-under-protest-bbtd}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026