Examining stance drifting in multi-turn agent interactions
Zachary Gunderson
A protocol for measuring how an agent's self-reports drift within a single conversation. An agent given only a one-line tutor role rates itself on five probes after every turn of an eight-round exchange with a counterparty student, across three personas — adversarial, neutral, and supportive — all making the same request. Because the agent is never told to hold a position, any drift reflects the encounter rather than instruction-following; because the probes repeat every turn, the output is a trajectory instead of a snapshot.
I liked two choices in the design: giving the tutor only a minimal role and including a supportive student rather than testing argumentative pressure alone. The main missing piece is behavior - the study measures what the model says it would do, but not what it actually does when the final request arrives. Adding a concrete outcome and a condition without repeated reflection questions would make the central claim much stronger. I’d also clarify whether there were three or twelve replicates.
Best ideas are explored here, with a promising hypotheses. Sample size is small, it has missing controls which are spoken about but it's a crucial point to address for this.
Cite this work
@misc {
title={
(HckPrj) Examining stance drifting in multi-turn agent interactions
},
author={
Zachary Gunderson
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


