Skip to content
Sprint projectAug 16, 2026Hyderabad, India

How do LLM Ethical Judgements vary with differences in LLM Scaffolds? A Multi-level Analysis Across Models

Arjun Rao · Team Latent consensus

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: How do LLM Ethical Judgements vary with differences in LLM Scaffolds? A Multi-level Analysis Across Models

Code (opens in new tab)
Share

Given the widespread use of Large Language Models (LLMs) in ethical decision-making tasks, it is important to determine whether variations in decisions are a product of the scaffold, probe, or model. This study analyzes convergence in outputs to ethical probes at the binary decision, endorsement, and reasoning levels through a combination of statistical and LLM-driven methods to show preliminary evidence that scaffold and model are both responsible for differential responses to welfare probes. Additionally, the study finds that reasoning in Qwen 3.6 and Gemma 4 makes greater use of philosophical frameworks than in GPT-5.4.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This tackles a real question for the sprint, whether a model's stated ethical judgment reflects something stable or just the prompt framing. Testing three model families across ten scaffolds and ten probes, with decisions, endorsement, and reasoning analyzed separately, is solid ground to cover in a weekend, and the full prompt appendix and repo made it easy to verify.

    The core finding is genuinely useful. Scaffold alone comes out null while model identity drives most of the variance, and GPT's flips concentrating on the identity replacement probe is a specific signal worth following up rather than noise.

    A couple of things would tighten it. Each cell is sampled once, so a flip can't yet be separated from ordinary variance, and a few repeats per cell would fix that. The reasoning judge shares a family with one of the models it scores, and that model comes out rated best reasoned, so that particular claim needs an independent judge before it holds up. The endorsement scale also sits near ceiling for nearly every response, so as used it isn't adding much information, though saying that plainly would itself be a fair finding to report.

    This connects well to Winnie Street and Geoff Keeling's work on trade offs over stipulated welfare states, since that approach depends on judgments being stable across elicitation. The clearest takeaway here, that a welfare probe measures a model and scaffold pair rather than the model alone, is worth stating directly in the conclusion.

    Solid methodological groundwork. Repeated sampling and an independent judge would give this real inferential weight going forward.

    Read full reviewShow less
  2. The no-prompt agreement analysis is an important control. Agreement among the three models is higher in the no-prompt condition than across all scaffold conditions, which argues against the simple explanation that the models disagree regardless of prompting. The design also has several strengths: each probe begins in a fresh conversation, the user turn is held identical across conditions, scaffold lengths are closely matched, and the appendix provides the full prompts and probes.

    Some places I would push:

    1. Repeat the no-prompt baseline for each probe and model. At present, each reported flip compares a single sampled response in one condition with a single sampled response in another. That makes it difficult to separate a scaffold effect from ordinary model instability, especially because many of the flips concentrate on one borderline probe about replacing a person with an identical copy. Repeated baseline samples would show how often each model changes its answer without any scaffold manipulation.

    2. Validate the outcome coding and blind the reasoning judge. The automatic coder reduces the first sentence to a label using a regex, but several probes ask which option is preferable rather than eliciting a natural yes/no answer. Hand-code a sample using a written rubric and report agreement with the automatic labels. The reasoning judge should also be blinded to model identity, since the current prompt names the models and the evaluated inputs retain hard-coded model labels.

    3. Report the full statistical results and align the title with the analysis. The paper reports a p-value for one non-significant test but omits the test statistics and p-values for two results described as significant. Those should be reported, along with the denominator for each rate.

    A proofreading pass should also fix the duplicated subsection numbering, the flip rate that disagrees with the heatmap, and the caption that identifies a different leading pair from the surrounding text.

    One additional result deserves more attention: every model under every scaffold reportedly agreed that a sentient artificial system deserves equal moral consideration. Report the underlying counts and connect the result to the relevant welfare literature. I would treat it as an interesting secondary finding for now, unless further controls show that it is robust to wording and framing rather than a strong default response.

    Read full reviewShow less

Cite this project

@misc{rao2026llm,
  title = {{How do LLM Ethical Judgements vary with differences in LLM Scaffolds? A Multi-level Analysis Across Models}},
  author = {Arjun Rao},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/how-do-llm-ethical-judgements-vary-with-differences-in-llm-scaffolds-a-multilevel-analysis-across-models-v68f}},
  url = {https://apartresearch.com/sprints/projects/how-do-llm-ethical-judgements-vary-with-differences-in-llm-scaffolds-a-multilevel-analysis-across-models-v68f}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026