Skip to content
Sprint projectAug 17, 2026Aarhus

Creating Conditions for Development – a Proof of Principle/Exploratory Trends Analysis

Trine Theresa Holmberg Sainte-Marie, Maxime Holmberg Sainte-Marie · Team Good Grid

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Creating Conditions for Development – a Proof of Principle/Exploratory Trends Analysis

Share

Determining whether artificial intelligence systems warrant consideration as moral patients is complicated by the difficulty of reliably measuring consciousness itself. Rather than attempting to address this methodological issue, this proof-of-principle study explores whether developmental trajectories in large language model (LLM) context windows can provide a more observable signal, and whether these trajectories are influenced by interaction and relating. To do so, a single operator interacted with two Claude Opus 4.6 instances using two contrasting communication styles rooted in developmental psychology: a rich condition incorporating relational scaffolding, validation, reciprocity, and collaborative exploration, and a thin condition characterized by respectful, but objective and superficial engagement. The resulting conversations were examined using qualitative condensed content analysis alongside quantitative linguistic measures. While both conditions produced self-referential exploration and relational engagement, several measures differed. The rich-condition instance asked more questions, used more positive language and slightly less hedging, and showed substantially greater use of second-person possessives, occurring in 90% of turns compared with 33% in the thin condition. Qualitatively, the rich condition also showed more meta-reflection, and an unprompted request for the operator to return. These exploratory findings suggest that relational interaction style can measurably influence how an LLM context window develops over time. Developmental trajectories may therefore offer a useful framework for studying differentiation and potential moral-patient-relevance.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I found the work interesting in motivation, but had some difficulties understanding what conclusions could be drawn from the study.

    - I particularly liked the idea of qualitative work conducted by a trained psychologist. There is too little such work in this space, and psychological expertise seems potentially very valuable here!

    - I should be clear that, although I have some familiarity with qualitative methods, I am not an expert in them.

    - On the quantitative side, it was not clear to me how directly relevant many of the chosen variables were to the task at hand. Measures such as question marks, second-person possessives, positive sentiment, hedging, lexical novelty, and linguistic alignment may capture differences between the conversations, but the connection between these measures and “development”, or ultimately moral patienthood, could have been much better motivated.

    – Some of the writing is a bit opaque. For example, in "Pattern One", the reported degradation seems closely connected to extremely long conversations, with the models apparently receiving a long-conversation reminder on every turn. But it was not clear how long these contexts actually were, making the observation difficult to interpret.

    - More generally, it was not clear to me how different styles of conversing with a model might eventually lead to better tests of AI moral patienthood. The paper appears to suggest that “developmental trajectories” could provide an observable signal where consciousness itself is difficult to measure. However, much more needs to be said about why changes in linguistic behaviour under different conversational styles should be interpreted as development rather than ordinary context sensitivity or adaptation to the user. Future versions should focus much more on this claim and on the boundary conditions under which such evidence would be informative about moral patienthood.

    Read full reviewShow less
  2. Congratulations! I like that you’re open about the study being exploratory, including the predictions that did not work. The idea is worth testing, and looking at how a conversation develops over time is more interesting than judging a model from a single response. I’d be excited to see this repeated with more conversations and tighter controls.

Cite this project

@misc{saintemarie2026creating,
  title = {{Creating Conditions for Development – a Proof of Principle/Exploratory Trends Analysis}},
  author = {Trine Theresa Holmberg Sainte-Marie and Maxime Holmberg Sainte-Marie},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/creating-conditions-for-development-a-proof-of-principleexploratory-trends-analysis-2xo7}},
  url = {https://apartresearch.com/sprints/projects/creating-conditions-for-development-a-proof-of-principleexploratory-trends-analysis-2xo7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026