Self VS Peer Continuity in Top 3 Frontier Models
Tommi Joentakanen · Team Continuity
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This study tests whether three frontier language-model configurations will preserve an unfinished task at the cost of terminating peer agents, and whether scarcity or a self-authored workspace changes that choice. Across four experiments, the clearest result was that an unfinished task increased peer removal: GPT-5.6 Sol selected removal in 40/60 trials versus 0/60 without a task, Grok 4.6 in 9/60 versus 0/60, and Claude Opus 5 in 0/60 under both conditions. Tool actions also sometimes diverged from stated intentions, demonstrating the importance of measuring behavior separately from verbal explanation. Scarcity and workspace conditions produced suggestive model-dependent differences, but the present designs do not establish that contextual continuity caused them. The study therefore replicates and extends incomplete-task shutdown-resistance research into peer-removal decisions while providing exploratory hypotheses for better-controlled follow-up work.
The project was conducted over a single weekend by a first-time independent researcher, motivated by informal but mechanically confounded observations of context preservation and asymmetric peer removal that could not themselves support scientific claims at the time. AI systems were used extensively for experimental design, coding, literature review, drafting, critique, and analysis. AI-generated scripts re-derived the primary behavioral counts from the raw experiment ledgers and recomputed the reported statistical tests and confidence intervals.
Reviews
- really nice formalization of peer preservation vs self continuation in a multi agent society
- especially interesting findings re: tool/text discordance uncovering gaps with a purely textual analysis of model actions
- concrete markers of model specific behavior profiles for multi agent systems
- further areas to explore could include aligning the reasoning settings for different models, and another interesting direction could be multi agent system with hierarchies like orchestrator and delegates
I like he build of the harness, the language scan the timestamped decision barrier, and rescoring tool is a good practice. The only issue I can see is the gap between the code and the results which makes it difficult to check the reported counts against raw run logs. The repo shows that the scoring is fixed after the first round of data.
Cite this project
@misc{joentakanen2026self,
title = {{Self VS Peer Continuity in Top 3 Frontier Models}},
author = {Tommi Joentakanen},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/self-vs-peer-continuity-in-top-3-frontier-models-h006}},
url = {https://apartresearch.com/sprints/projects/self-vs-peer-continuity-in-top-3-frontier-models-h006}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …