Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior
Cole Alexander Niblett, Sofiia Lobanova
The contents of a model's context window play an important role in determining its behavior, a fact frequently exploited in jailbreaking. Under the persona selection model, this is conceptualized as an inference-time indexing of a distribution over model personas, learned during pre-training and re-weighted during post-training. The factors that influence persona indexing are a critical issue for alignment, determining the conditions under which safety training generalizes to real-world deployment scenarios. One specific category of context element that may serve as a persona-indexing feature is tool affordances. Prior work by Lynch et al. (2025) and Kutasov et al. (2026) suggests that the presence of tool affordances in deployment contexts may drive misalignment in models trained on plain-chat training environments. We ask whether a specific failure mode, task refusal, can be causally influenced by the presence of various tool affordances, finding that tool presence induces refusal from a zero baseline in four of eleven models, most strongly under the most inert tool, and shifts how models describe themselves, toward a more instrumental, in-service self-account.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior
},
author={
Cole Alexander Niblett, Sofiia Lobanova
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


