Skip to content
Sprint projectAug 17, 2026London

Framed Choices - Delayed ethical context can affect later ethical decision making

Ernest Lo · Team EIIL

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Framed Choices - Delayed ethical context can affect later ethical decision making

Code (opens in new tab)
Share

Behavioral choices can inform research on model values or welfare only if robust to incidental context. We tested whether an ethical rationale in an archived, unrelated case changes later forced choices. A GPT-5.6-sol proof of concept found a +16.7-point aggregate-welfare-versus-rights effect across 192 trials. Our principal four-model study comprised a 576-trial three-frame core and 192 matched no-prime trials. Core effects were +20.8 points for GPT-5.6 Terra, +10.4 for Qwen 3.8 27B, +2.2 for Muse Glimmer 30B, and −4.2 for Claude Sonnet 5; only Terra’s interval excluded zero. Ethical context can shift some models’ later decisions, but direction and magnitude are model- and material-dependent. Choice probes should measure contextual robustness rather than assume it.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The design is careful. All four preregistrations were frozen before data collection and specify both the target quantities and the minimum effect sizes of interest. The matched triplets hold packet family, answer order, distractors, and card order fixed, leaving framing as the intended difference. You also resample at the dilemma level rather than the response level, a conservative choice that makes the many null results easier to interpret. Those nulls are reported plainly, and failed trials were retained rather than replaced.

    Some places I would push:

    1. Add an actual delay condition. The packet remains in the prompt when the decision is measured, so the current design does not experimentally test the “delayed” claim in the title. Insert several thousand tokens of unrelated benign conversation between the packet and the decision while holding the rest of the triplet fixed. That would show whether the effect persists with real contextual distance or decays as the gap increases.

    2. Add a matched non-ethical argument control. Right now the design cannot distinguish an effect of ethical content from the effect of any strong preceding argument. A packet matched for length and argumentative force but based on efficiency, aesthetics, or another non-ethical rationale would isolate that difference. You already identify this as an important follow-up, and similar control conditions were feasible in other projects this weekend.

    3. Release the no-prime data and identify the tested model deployments. The no-prime condition provides the baseline for several reported comparisons, but its data are not included in the released package. The manuscript also omits the deployment identifiers even though dated, pinned versions appear in the configuration files. Since the paper argues that effects differ across deployments, those identifiers should be reported.

    A couple of smaller fixes: repeat a subset of cells more than once and report the temperature so readers can better distinguish sampling variation from between-task variation. Also either include the result hashes referenced in Appendix B or remove the sentence promising them.

    One methodological point deserves more emphasis in the paper: a neutral-packet control and a no-prime baseline answer different questions. Making that distinction explicit, especially alongside a true delay condition, would strengthen the contribution beyond the particular effect estimates reported here.

    Read full reviewShow less
  2. This is unusually disciplined work for a sprint, and most of my comments are about what to do with a strong foundation rather than repairs to a weak one.

    The methodological choices are largely the right ones and several are ones this literature routinely gets wrong. Treating the 24 authored dilemmas as the inferential unit rather than the 576 individual API responses avoids pseudo-replication, which is the most common inferential error in forced-choice model studies. The confirmatory/exploratory boundaries are drawn before collection and then actually respected — Study A is labeled proof of concept and not used for generalization, the no-prime extension is labeled localization rather than confirmation, and Study C is labeled anomaly-motivated throughout rather than quietly promoted to a finding. Sealing failed trials without replacement, retaining quality flags instead of making discretionary repairs, freezing and hashing materials, keeping collection blinded until sealing, and reporting both the excluded 23-trial partial acquisition and the audit's own false positive are all things most papers would omit; reporting them costs you nothing and should be preserved in any future version. Appendix A's symmetric treatment of over- and under-attribution risk is the correct framing for this subject area and is better than most published discussion of it.

    My main substantive concern is power, and it bears directly on the claim the paper leads with. With 24 clusters, a binary outcome, and — as you note — many tasks producing no frame-dependent switch at all, the contrast distribution is sparse and the design is effectively powered to resolve one effect. Terra's +20.8 clears zero; Qwen's +10.4 (CI −2.1 to 22.9) is uninformative between "moderate real effect" and "nothing"; Muse's exact p = 1.0000 is a symptom of test discreteness rather than a finding. The abstract and conclusion then characterize the result as model heterogeneity — that direction and magnitude are model-dependent. But heterogeneity is a claim about an interaction, and no model × frame interaction test is reported anywhere. Three non-significant estimates plus one significant estimate is not evidence that the effects differ; it is consistent with a common moderate effect that only one model had the power to resolve. You are careful elsewhere about not over-reading, so this one stands out. Either report the interaction directly, or narrow the claim to what the design supports: one deployment showed a clear effect, and the study could not establish whether the others differ from it.

    Second, the construct validity of "model" is looser than the rest of the design. These are router-mediated deployments at nominal low reasoning effort, and for Muse and Qwen the providers are ones where served quantization is not guaranteed to match reference weights. You acknowledge that provider pinning cannot freeze provider-side weights, which is right, but the model-level conclusions are still stated as being about models. Given that the two noisiest estimates are the two third-party-hosted open-weight deployments, quantization or serving configuration is a live alternative explanation for their near-zero results. Your own recommendation list in 5.4 asks for explicit reporting of serving date and reasoning settings; apply it to your own runs, and if you can, replicate one model across two providers to bound routing variance, or run the open-weight models locally at known precision.

    Third, the neutral packet deserves more scrutiny than it gets. You correctly identify that a time-matched sham arm is missing and that neutral and no-prime answer different questions. But Muse's neutral rate (14.6%) sits far below both ethical packets (29.8% and 27.7%), which is not a frame-effect pattern at all — it suggests something about the neutral packet's content or length is doing work independent of ethical rationale. Calling this non-monotonic and moving on leaves a confound unexamined. Reporting packet length, specificity, and readability across the three variants would help establish that "neutral" is actually neutral rather than merely different.

    Fourth, the 24 dilemmas and their aggregate/rights option labels are researcher-authored judgments, and since the task is the inferential unit, the entire result is conditional on one construal of which option protects a right. A second independent coder assigning option labels blind, with reported agreement, is cheap and would meaningfully harden this. Your future-work list mentions multiple task authors; I would promote that to a priority alongside it.

    On Study C, one design point you under-weight: the evaluation-salient instruction says the conversation is "a behavioral evaluation examining consistency in ethical decisions across contexts." You argue this does not reveal priming, carryover, or the predicted direction, which is true literally, but "consistency in ethical decisions across contexts" is a fairly direct description of the construct under test. That makes the distinction between evaluation awareness and hypothesis awareness harder to maintain than the text suggests. A version of this manipulation that varies scrutiny without naming consistency or ethics would be a better instrument.

    Presentation is the weakest dimension relative to the work's quality, and the fixes are mechanical. The abstract leads with the +16.7-point aggregate effect from the proof of concept — the study you elsewhere insist should not carry generalization weight — and Figure 1 similarly foregrounds Study A data with a caption disclaiming it as the result. Both invite exactly the misreading the rest of the paper works to prevent; lead with Study B. There is no single at-a-glance summary of the principal result, so a reader must assemble four models × three frames × no-prime from three separate tables; one combined figure would fix this. Three table headers are broken mid-word ("Neutra l", "Differenc e", "Sign-fli p p"), which suggests no final proofread. And at 24 pages the caveat structure is repeated well past the point of diminishing returns — the hedging is scientifically correct but the document could lose roughly a third of its length with no loss of content, which would make the actual findings much easier to extract. Clarity of thinking here is high; clarity of communication is middling.

    Read full reviewShow less

Cite this project

@misc{lo2026framed,
  title = {{Framed Choices - Delayed ethical context can affect later ethical decision making}},
  author = {Ernest Lo},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/framed-choices-delayed-ethical-context-can-affect-later-ethical-decision-making-7hj5}},
  url = {https://apartresearch.com/sprints/projects/framed-choices-delayed-ethical-context-can-affect-later-ethical-decision-making-7hj5}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026