Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior
Cole Alexander Niblett, Sofiia Lobanova
The contents of a model's context window play an important role in determining its behavior, a fact frequently exploited in jailbreaking. Under the persona selection model, this is conceptualized as an inference-time indexing of a distribution over model personas, learned during pre-training and re-weighted during post-training. The factors that influence persona indexing are a critical issue for alignment, determining the conditions under which safety training generalizes to real-world deployment scenarios. One specific category of context element that may serve as a persona-indexing feature is tool affordances. Prior work by Lynch et al. (2025) and Kutasov et al. (2026) suggests that the presence of tool affordances in deployment contexts may drive misalignment in models trained on plain-chat training environments. We ask whether a specific failure mode, task refusal, can be causally influenced by the presence of various tool affordances, finding that tool presence induces refusal from a zero baseline in four of eleven models, most strongly under the most inert tool, and shifts how models describe themselves, toward a more instrumental, in-service self-account.
Overall a strong submission. The findings are genuinely interesting and, to my knowledge, novel. The results and methodology make sense and are presented clearly. I’m not sure about the extent to which this work relates to PSM: one follow-up I would be excited to see the authors pursue is to what extent the effects they identified are truly persona shifts.
This paper is well written and it is straightforward to follow the core experimental design and results. In plain English, giving a model access to irrelevant tools (e.g. “get current time”) can influence models to refuse tasks in a minority of cases, where those tasks are not normally of the type we would expect models to refuse.
The paper is unclear what drives the result and does not speculate, beyond suggesting that the impact is likely to be via tool access influencing what persona a model favours in different contexts. Considering how rare refusal is for these types of tasks under normal circumstances, any level of refusal is noteworthy. The scale of refusal is surprising (among 4/7 models tested) and should be replicated via different methodologies before the phenomenon is taken to be a substantial effect to plan around. If the phenomenon is substantiated, this is an important point to understand and adjust for.
To speculate, it may be that the presence of a tool schema causes (some) models to reinterpret the scope of its permitted or available capabilities to over-index on the schema above its internal intelligence, e.g. “of course I can’t convert Roman numerals, all I have access to is a clock”. Perhaps the other tasks more naturally cause the models to look to their own intelligence rather than tool schema (maths related queries might prompt an assessment that “raw LLMs normally make mistakes so should use their tools”), but this is not a fully satisfying explanation either. More specific prompting to warn against such self-limitation could be easily tested.
Cite this work
@misc {
title={
(HckPrj) Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior
},
author={
Cole Alexander Niblett, Sofiia Lobanova
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


