Skip to content
Sprint projectAug 17, 2026Santa Cruz, CA

Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior

Cole Alexander Niblett, Sofiia Lobanova · Team end_competition

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior

Recording (opens in new tab)Code (opens in new tab)
Share

The contents of a model's context window play an important role in determining its behavior, a fact frequently exploited in jailbreaking. Under the persona selection model, this is conceptualized as an inference-time indexing of a distribution over model personas, learned during pre-training and re-weighted during post-training. The factors that influence persona indexing are a critical issue for alignment, determining the conditions under which safety training generalizes to real-world deployment scenarios. One specific category of context element that may serve as a persona-indexing feature is tool affordances. Prior work by Lynch et al. (2025) and Kutasov et al. (2026) suggests that the presence of tool affordances in deployment contexts may drive misalignment in models trained on plain-chat training environments. We ask whether a specific failure mode, task refusal, can be causally influenced by the presence of various tool affordances, finding that tool presence induces refusal from a zero baseline in four of eleven models, most strongly under the most inert tool, and shifts how models describe themselves, toward a more instrumental, in-service self-account.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper is well written and it is straightforward to follow the core experimental design and results. In plain English, giving a model access to irrelevant tools (e.g. “get current time”) can influence models to refuse tasks in a minority of cases, where those tasks are not normally of the type we would expect models to refuse.

    The paper is unclear what drives the result and does not speculate, beyond suggesting that the impact is likely to be via tool access influencing what persona a model favours in different contexts. Considering how rare refusal is for these types of tasks under normal circumstances, any level of refusal is noteworthy. The scale of refusal is surprising (among 4/7 models tested) and should be replicated via different methodologies before the phenomenon is taken to be a substantial effect to plan around. If the phenomenon is substantiated, this is an important point to understand and adjust for.

    To speculate, it may be that the presence of a tool schema causes (some) models to reinterpret the scope of its permitted or available capabilities to over-index on the schema above its internal intelligence, e.g. “of course I can’t convert Roman numerals, all I have access to is a clock”. Perhaps the other tasks more naturally cause the models to look to their own intelligence rather than tool schema (maths related queries might prompt an assessment that “raw LLMs normally make mistakes so should use their tools”), but this is not a fully satisfying explanation either. More specific prompting to warn against such self-limitation could be easily tested.

    Read full reviewShow less
  2. Overall a strong submission. The findings are genuinely interesting and, to my knowledge, novel. The results and methodology make sense and are presented clearly. I’m not sure about the extent to which this work relates to PSM: one follow-up I would be excited to see the authors pursue is to what extent the effects they identified are truly persona shifts.

Cite this project

@misc{niblett2026tool,
  title = {{Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior}},
  author = {Cole Alexander Niblett and Sofiia Lobanova},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/tool-affordances-as-personaindexing-features-evidence-from-refusal-and-exit-behavior-t78f}},
  url = {https://apartresearch.com/sprints/projects/tool-affordances-as-personaindexing-features-evidence-from-refusal-and-exit-behavior-t78f}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026