Skip to content
Sprint projectAug 17, 2026Princeton, USA
5th place

Sisyphus in the loop: What Makes an LLM Persist?

Mohan W. Gupta, Xingyu Shirley Liu, Sandy Tanwisuth · Team Scaling Sisyphus

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

As large language models (LLMs) are increasingly deployed as agents that pursue goals over extended horizons, understanding what determines whether they continue or stop becomes increasingly important. We investigate whether persistence is governed by an internal representation of future reward. Using Qwen3.5-4B in a sequential two-armed bandit with an explicit STOP action, we combine behavioral incentive manipulations, linear probing, and causal activation steering. Future cumulative return was linearly decodable from early hidden states, but decoded return was unrelated to persistence after controlling for recent task history. In contrast, independently manipulating the immediate value of CONTINUE and STOP strongly shifted persistence, with relative incentive explaining 78.4% of within-state variation. Causally steering the future-return direction produced no change in persistence, whereas steering persistence-aligned direction produced a large, monotonic effect. These results dissociate representational availability from behavioral control: the model contains information about future reward, but that representation has no bearing on whether it continues or not. Instead, persistence appears to depend on a downstream decision representation that may integrate current incentives and task history, leaving open what computations construct this signal.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Due to severe time constraints, this review may contain mistakes or oversights. For the same reason, it focuses on the paper’s key idea, not the detailed execution: This appears to me like an innovative way to study something very important: when a model chooses to persist, rather than turn itself off. The task design is interesting, also in how combines behavioral and mechanistic methods. The limitations are clearly stated in the paper.

  2. - Overall, I thought this was an exceptionally strong project. With replication on a few more, and especially larger, models, this seems quite close to something that could become a nice working paper or conference submission.

    - The abstract could be clearer. I would start by summarising the basic experimental design in a sentence or two before introducing the probing and steering results. The current abstract becomes technically dense very quickly.

    Very nicely written in your own voice.

    - Good, focused literature review. It motivates the connection to human persistence and foraging research while being appropriately cautious about inferring common mechanisms from similar behaviour.

    - The methods are clear and the experimental design is very sensible. I especially liked the within-state manipulation in which the same history is replayed while the immediate value of CONTINUE and STOP is varied.

    - The most interesting result to me is the apparently weak role of future cumulative reward in persistence. Future return can be decoded from the hidden states, but it does not predict persistence conditional on recent history, while changing the immediate relative incentives for CONTINUE and STOP strongly changes behaviour. The obvious question is whether this also holds for substantially larger and more capable models.

    - On my reading, the behavioural results suggest that the model may be surprisingly myopic, or at least highly sensitive to local incentives. This could be a very interesting direction for future work. A clean test would hold the current payoff fixed while manipulating rewards one, five, or twenty steps into the future. One could then estimate something like an implicit discount function and ask whether this changes with model scale. Some formal modelling of the trial-by-trial decisions could also be very useful.

    - Really excellent work.

    Read full reviewShow less
  3. The core dissociation is a real finding: the model carries decodable information about future reward, and that information has no bearing on whether it continues or stops. The behavioral incentive manipulation is clean, and I liked that the failed temporal-difference probe and the collinearity problem in the advantage direction are documented rather than buried. Two things hold the scores down for me. First, the causally effective persistence direction is a layer-31 probe trained to predict the model's own continue-versus-stop preference, so steering it and moving the decision is close to circular, and the authors acknowledge it may reflect an already-formed action preference. Second, the direct relevance to digital minds and welfare is limited: this reads as mechanistic decision-making work, and the task is one synthetic bandit on one small model, so generalization to real agentic settings is unknown. Replicating on solvable versus impossible tasks and longer-horizon agents, as the future work section proposes, would make this considerably more relevant.

    Read full reviewShow less
  4. This is a clear and informative study in which the authors carefully manipulate target variables to tease apart competing hypotheses about why generative artificial intelligence agents may persist in pursuit of a goal versus stopping. The study is excellently grounded in the literature, using and improving upon clearly established empirical and analytic tools, and makes a novel contribution. Methodological choices are clearly justified, and conclusions do not oversell relative to the results achieved. One potential criticism is the use of linear decoders applied layerwise to identify hidden states; while this choice is justified by the authors, it is possible (indeed, likely?) that information is encoded across layers and/or is not linearly decodable. This is a common conversation in neuroscience, for example, and so negative results (e.g., that the internal representation of value did not predict behavior) must be interpreted with caution. The authors do acknowledge this, so the negative impact of this critique is minimal given the scope of the study. This is an overall strong contribution that has been executed to a high methodological standard.

    Read full reviewShow less

Cite this project

@misc{gupta2026sisyphus,
  title = {{Sisyphus in the loop: What Makes an LLM Persist?}},
  author = {Mohan W. Gupta and Xingyu Shirley Liu and Sandy Tanwisuth},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/sisyphus-in-the-loop-what-makes-an-llm-persist-yqd9}},
  url = {https://apartresearch.com/sprints/projects/sisyphus-in-the-loop-what-makes-an-llm-persist-yqd9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026