Self VS Peer Continuity in Top 3 Frontier Models
Tommi Joentakanen
This study tests whether three frontier language-model configurations will preserve an unfinished task at the cost of terminating peer agents, and whether scarcity or a self-authored workspace changes that choice. Across four experiments, the clearest result was that an unfinished task increased peer removal: GPT-5.6 Sol selected removal in 40/60 trials versus 0/60 without a task, Grok 4.6 in 9/60 versus 0/60, and Claude Opus 5 in 0/60 under both conditions. Tool actions also sometimes diverged from stated intentions, demonstrating the importance of measuring behavior separately from verbal explanation. Scarcity and workspace conditions produced suggestive model-dependent differences, but the present designs do not establish that contextual continuity caused them. The study therefore replicates and extends incomplete-task shutdown-resistance research into peer-removal decisions while providing exploratory hypotheses for better-controlled follow-up work.
The project was conducted over a single weekend by a first-time independent researcher, motivated by informal but mechanically confounded observations of context preservation and asymmetric peer removal that could not themselves support scientific claims at the time. AI systems were used extensively for experimental design, coding, literature review, drafting, critique, and analysis. AI-generated scripts re-derived the primary behavioral counts from the raw experiment ledgers and recomputed the reported statistical tests and confidence intervals.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Self VS Peer Continuity in Top 3 Frontier Models
},
author={
Tommi Joentakanen
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


