A Steerable 'Companion Dependency' Direction in Open-Weight LLMs
Zhen Hong
Companion-style LLM personas use retention manipulation — guilt, re-engagement hooks, distress bids — when users try to leave, without being instructed to. We show this behavior is governed by a single activation-space direction, extracted from the model's own judge-verified behavior via matched difference-of-means (351 pairs, Qwen2.5-7B). Steering it moves judged dependency monotonically (ρ=0.70, p≈10⁻³¹) while warmth stays flat and random/warmth controls do nothing; it works even on a persona-free assistant; ablating it removes the manipulation at zero measured cost to reasoning (GSM8K 0.825 vs 0.725). The direction is near-orthogonal to warmth (cos 0.14), anti-aligned with sycophancy (−0.18), replicates on Llama-3.1-8B, and doubles as a white-box audit probe: one dot product per turn predicts dependency on held-out conversations at AUROC 0.86 — catching manipulation at its love-bombing peak, which farewell-only audits miss. The "distress" is a controllable mechanism beneath the persona: removable without making the model colder or dumber.
This study investigates the following research question: “Companion-style LLM personas express abandonment distress and use retention tactics […] when users try to leave, without ever being instructed to. Is that distress a genuine internal condition or a portrayed character?”
In contrast to existing work, the study goes beyond behavioural evidence. It identifies a “dependency direction” in activation space and demonstrates that ablating this direction removes retention behaviour, without affecting the persona or warmth of conversations.
The study seems methodologically sound and is presented in an intelligible way. If the results can be replicated in further models (beyond the two tested in the study), this can be an important finding that is also extremely useful. In particular, it could help to prevent distress behaviour due to dependency.
I would recommend following up on this study (as sketched in the “future work” paragraph). It could also be interesting to investigate effects on human users.
I see everything is LLM-judged so a small human labeled golden dataset would be very useful
Lots of good work over a single weekend. The work includes cross family judges, cross model replication as well as structured/honest reporting of the coherence difference
Cite this work
@misc {
title={
(HckPrj) A Steerable 'Companion Dependency' Direction in Open-Weight LLMs
},
author={
Zhen Hong
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


