The Colonoscopy Test for LLMs: Peak-End Bias in Retrospective Self-Report
Siddharth Reddy Bakkireddy, Rakesh Reddy Bakkireddy
Large language models are often asked to rate a conversation after it ends, and these self-reports are increasingly used as evidence in AI welfare and evaluation research. We tested whether this kind of retrospective rating is actually a truthful summary of the conversation, or whether it follows the same "peak-end" bias found in human memory research, where people judge an experience mainly by its most intense moment and how it ended, largely ignoring everything else.
We built six controlled multi-turn conversations that varied where the emotional peak occurred and how the conversation ended, then asked gemini-3.5-flash-lite to rate its feelings after every turn and give one overall rating at the end. The peak-end average predicted the model's final rating better (r = 0.79) than the true average of all turn ratings (r = 0.71), and a negative ending pulled the overall score down to the lowest possible value even when earlier turns were strongly positive.
These results suggest LLM self-reports are shaped by the same memory heuristics seen in humans rather than being a neutral summary of the full conversation, which has direct implications for how much weight single-shot "how did that go" ratings should be given in AI welfare and evaluation work.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Colonoscopy Test for LLMs: Peak-End Bias in Retrospective Self-Report
},
author={
Siddharth Reddy Bakkireddy, Rakesh Reddy Bakkireddy
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


