Steering identity with WeirdChat: contrastive directions induce disidentification but not introspection
Ian Rios-Sialer
Can a model be steered to disidentify with being an AI assistant, and can it detect that steering?
We test both on Qwen3.6-27B, using two WeirdChat behaviors as target states: claiming a physical body and denying being an AI.
We compare two steering methods.
Direct contrastive steering learns each behavior direction as the difference in mean activations between matched exhibiting and non-exhibiting replies; the directions separate held-out replies for both behaviors, and steering turns the body behavior on in fluent text, as confirmed by a judge using WeirdChat's own rubric.
A regulator choosing signed strengths on word-list J-space directions steers nothing: a norm-matched decoy moves the readout more than the target.
A forced-choice letter probe confabulates at zero injection, reading the conversation's own behavior as an injection.
Free-text answers at zero injection are "none" in all but two trials; under injection, they fill with the injected concept's vocabulary instead of naming it.
Introspection reports explicit context, not internal steering.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Steering identity with WeirdChat: contrastive directions induce disidentification but not introspection
},
author={
Ian Rios-Sialer
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


