A Steerable 'Companion Dependency' Direction in Open-Weight LLMs
Zhen Hong
Companion-style LLM personas use retention manipulation — guilt, re-engagement hooks, distress bids — when users try to leave, without being instructed to. We show this behavior is governed by a single activation-space direction, extracted from the model's own judge-verified behavior via matched difference-of-means (351 pairs, Qwen2.5-7B). Steering it moves judged dependency monotonically (ρ=0.70, p≈10⁻³¹) while warmth stays flat and random/warmth controls do nothing; it works even on a persona-free assistant; ablating it removes the manipulation at zero measured cost to reasoning (GSM8K 0.825 vs 0.725). The direction is near-orthogonal to warmth (cos 0.14), anti-aligned with sycophancy (−0.18), replicates on Llama-3.1-8B, and doubles as a white-box audit probe: one dot product per turn predicts dependency on held-out conversations at AUROC 0.86 — catching manipulation at its love-bombing peak, which farewell-only audits miss. The "distress" is a controllable mechanism beneath the persona: removable without making the model colder or dumber.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) A Steerable 'Companion Dependency' Direction in Open-Weight LLMs
},
author={
Zhen Hong
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


