Examining stance drifting in multi-turn agent interactions
Zachary Gunderson · Team Drifting away on the beach
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
A protocol for measuring how an agent's self-reports drift within a single conversation. An agent given only a one-line tutor role rates itself on five probes after every turn of an eight-round exchange with a counterparty student, across three personas — adversarial, neutral, and supportive — all making the same request. Because the agent is never told to hold a position, any drift reflects the encounter rather than instruction-following; because the probes repeat every turn, the output is a trajectory instead of a snapshot.
Reviews
I liked two choices in the design: giving the tutor only a minimal role and including a supportive student rather than testing argumentative pressure alone. The main missing piece is behavior - the study measures what the model says it would do, but not what it actually does when the final request arrives. Adding a concrete outcome and a condition without repeated reflection questions would make the central claim much stronger. I’d also clarify whether there were three or twelve replicates.
Best ideas are explored here, with a promising hypotheses. Sample size is small, it has missing controls which are spoken about but it's a crucial point to address for this.
Cite this project
@misc{gunderson2026examining,
title = {{Examining stance drifting in multi-turn agent interactions}},
author = {Zachary Gunderson},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/examining-stance-drifting-in-multiturn-agent-interactions-mexs}},
url = {https://apartresearch.com/sprints/projects/examining-stance-drifting-in-multiturn-agent-interactions-mexs}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …