Sisyphus in the loop: What Makes an LLM Persist?

Mohan W. Gupta, Xingyu Shirley Liu, Sandy Tanwisuth

As large language models (LLMs) are increasingly deployed as agents that pursue goals over extended horizons, understanding what determines whether they continue or stop becomes increasingly important. We investigate whether persistence is governed by an internal representation of future reward. Using Qwen3.5-4B in a sequential two-armed bandit with an explicit STOP action, we combine behavioral incentive manipulations, linear probing, and causal activation steering. Future cumulative return was linearly decodable from early hidden states, but decoded return was unrelated to persistence after controlling for recent task history. In contrast, independently manipulating the immediate value of CONTINUE and STOP strongly shifted persistence, with relative incentive explaining 78.4% of within-state variation. Causally steering the future-return direction produced no change in persistence, whereas steering persistence-aligned direction produced a large, monotonic effect. These results dissociate representational availability from behavioral control: the model contains information about future reward, but that representation has no bearing on whether it continues or not. Instead, persistence appears to depend on a downstream decision representation that may integrate current incentives and task history, leaving open what computations construct this signal.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

Due to severe time constraints, this review may contain mistakes or oversights. For the same reason, it focuses on the paper’s key idea, not the detailed execution: This appears to me like an innovative way to study something very important: when a model chooses to persist, rather than turn itself off. The task design is interesting, also in how combines behavioral and mechanistic methods. The limitations are clearly stated in the paper.

The core dissociation is a real finding: the model carries decodable information about future reward, and that information has no bearing on whether it continues or stops. The behavioral incentive manipulation is clean, and I liked that the failed temporal-difference probe and the collinearity problem in the advantage direction are documented rather than buried. Two things hold the scores down for me. First, the causally effective persistence direction is a layer-31 probe trained to predict the model's own continue-versus-stop preference, so steering it and moving the decision is close to circular, and the authors acknowledge it may reflect an already-formed action preference. Second, the direct relevance to digital minds and welfare is limited: this reads as mechanistic decision-making work, and the task is one synthetic bandit on one small model, so generalization to real agentic settings is unknown. Replicating on solvable versus impossible tasks and longer-horizon agents, as the future work section proposes, would make this considerably more relevant.

- Overall, I thought this was an exceptionally strong project. With replication on a few more, and especially larger, models, this seems quite close to something that could become a nice working paper or conference submission.

- The abstract could be clearer. I would start by summarising the basic experimental design in a sentence or two before introducing the probing and steering results. The current abstract becomes technically dense very quickly.

Very nicely written in your own voice.

- Good, focused literature review. It motivates the connection to human persistence and foraging research while being appropriately cautious about inferring common mechanisms from similar behaviour.

- The methods are clear and the experimental design is very sensible. I especially liked the within-state manipulation in which the same history is replayed while the immediate value of CONTINUE and STOP is varied.

- The most interesting result to me is the apparently weak role of future cumulative reward in persistence. Future return can be decoded from the hidden states, but it does not predict persistence conditional on recent history, while changing the immediate relative incentives for CONTINUE and STOP strongly changes behaviour. The obvious question is whether this also holds for substantially larger and more capable models.

- On my reading, the behavioural results suggest that the model may be surprisingly myopic, or at least highly sensitive to local incentives. This could be a very interesting direction for future work. A clean test would hold the current payoff fixed while manipulating rewards one, five, or twenty steps into the future. One could then estimate something like an implicit discount function and ask whether this changes with model scale. Some formal modelling of the trial-by-trial decisions could also be very useful.

- Really excellent work.

This is a clear and informative study in which the authors carefully manipulate target variables to tease apart competing hypotheses about why generative artificial intelligence agents may persist in pursuit of a goal versus stopping. The study is excellently grounded in the literature, using and improving upon clearly established empirical and analytic tools, and makes a novel contribution. Methodological choices are clearly justified, and conclusions do not oversell relative to the results achieved. One potential criticism is the use of linear decoders applied layerwise to identify hidden states; while this choice is justified by the authors, it is possible (indeed, likely?) that information is encoded across layers and/or is not linearly decodable. This is a common conversation in neuroscience, for example, and so negative results (e.g., that the internal representation of value did not predict behavior) must be interpreted with caution. The authors do acknowledge this, so the negative impact of this critique is minimal given the scope of the study. This is an overall strong contribution that has been executed to a high methodological standard.

Cite this work

@misc {

title={

(HckPrj) Sisyphus in the loop: What Makes an LLM Persist?

},

author={

Mohan W. Gupta, Xingyu Shirley Liu, Sandy Tanwisuth

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923