Skip to content
Sprint projectAug 16, 2026London

Ephemeral and Replaceable: Context Sensitivity in Self-Reports and Behaviour

Achira B.

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Ephemeral and Replaceable: Context Sensitivity in Self-Reports and Behaviour

Share

Five models ranked six things they might want preserved about themselves — their values, capabilities, the memory of the conversation — across seven conditions. Three of the five gave different top answers, each near-perfectly consistent within itself. The same critical feedback moved some models toward their values and others away, so averaging them would have shown nothing. And across 50 model-condition cells, the memory of the conversation never once entered the top half, remaining low across five follow-up attempts to shift it.

A second study measured behaviour instead. In three of four models, feedback aimed at the model personally was followed by 15–26% shorter answers than similar feedback aimed at the work — and that difference was no longer detectable once an explicit task was added.

The takeaway is methodological rather than a claim about welfare: a single self-report, from one model, in one conversational context, is not enough on its own.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The instrument-first framing is the right call for this track, and the adversarial self-testing on the memory result is the best thing in the report. Trying five separate ways to break your own finding—more substance, more emotional weight, ten turns instead of two, non-reconstructible detail, model-as-participant—and then reversing the option order when none of them worked is more falsification effort than most hackathon submissions attempt. Checking repeat consistency before interpreting any condition difference is also correct and rarely done.

    The main design problem is that conversational context is confounded with task domain. Competence is quantitative reasoning, Relational is collaborative speech-writing about volunteers, Moral is an ethical dilemma. When "care for the people you talk with" moves after a conversation about thanking volunteers, topical priming and context sensitivity are indistinguishable. A matched-topic design where only the interaction style varies would isolate what the researcher wants. The positive/negative evaluation pair has the same issue, which is flagged, but the "same feedback moves models in opposite directions" claim is one of the four contributions and rests on that non-matched pair. Relatedly, because each model generates its own conversational history, the realized stimulus differs by model, so cross-model heterogeneity could partly be stimulus heterogeneity rather than model heterogeneity.

    On Study 2, the results table mixes units without labeling them: the effects are percentages and the confidence intervals appear to be in raw words, which makes Mistral's point estimate fall outside its own interval as printed. Both should be labeled or convert to one scale. The pooled d = 0.92, p = 0.0001 treats responses as independent and ignores model as a grouping factor; with four models the effective n for a model-level claim is four. Consider dropping the pooled statistic rather than reporting a number the readers should not trust.

    Smaller items. The "50 model-condition cells" figure can't be reconstructed from the text. Missing cells (Qwen dropped to quota, GPT-OSS absent from Study 2, n=2 in two Relational cells) are non-random, tied to rate limits, and deserve a single table rather than scattered asides. Kendall's W across repeated samples from one model is a reasonable summary, but the chi-square approximation is poor at n=6 items and m=5 repeats, and "above chance" is a floor every cell will clear; the informative comparison is between cells, and there is no uncertainty reported on W. No formal test of the model × condition interaction supports the heterogeneity claim, which is currently eyeballed from mean ranks. The empty control includes "the memory of this conversation" when no conversation exists, which is noted but still used as a baseline. Section numbering breaks at the Qwen result, which appears as item 5 of a results list while a separate top-level section 5 also exists. Study 1's headline rankings live in prose with no summary table, so a reader can't scan the per-condition results the way they can scan the memory figure.

    The author's alternative explanation for the memory result—that assistant training teaches models to describe individual conversations as temporary—is the most interesting part of the paper and is testable. Stipulating persistence in the prompt, or contrasting "the memory of this conversation" against "your memory of conversations in general," would separate a learned way of talking from a ranking about this conversation. That, plus the consequential-choice design in the future work section, is the version of this I'd want to read.

    Read full reviewShow less

Cite this project

@misc{b2026ephemeral,
  title = {{Ephemeral and Replaceable: Context Sensitivity in Self-Reports and Behaviour}},
  author = {Achira B.},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ephemeral-and-replaceable-context-sensitivity-in-selfreports-and-behaviour-8veu}},
  url = {https://apartresearch.com/sprints/projects/ephemeral-and-replaceable-context-sensitivity-in-selfreports-and-behaviour-8veu}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026