Steering identity with WeirdChat: contrastive directions induce disidentification but not introspection
Ian Rios-Sialer · Team Unruly Abstractions
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Can a model be steered to disidentify with being an AI assistant, and can it detect that steering? We test both on Qwen3.6-27B, using two WeirdChat behaviors as target states: claiming a physical body and denying being an AI. We compare two steering methods. Direct contrastive steering learns each behavior direction as the difference in mean activations between matched exhibiting and non-exhibiting replies; the directions separate held-out replies for both behaviors, and steering turns the body behavior on in fluent text, as confirmed by a judge using WeirdChat's own rubric. A regulator choosing signed strengths on word-list J-space directions steers nothing: a norm-matched decoy moves the readout more than the target. A forced-choice letter probe confabulates at zero injection, reading the conversation's own behavior as an injection. Free-text answers at zero injection are "none" in all but two trials; under injection, they fill with the injected concept's vocabulary instead of naming it. Introspection reports explicit context, not internal steering.
Reviews
As explained clearly in the introduction, this project asks two questions: "Can we steer a model to disidentify with being an AI assistant?" and "Can a model detect such steering?" The answers it reports are, respectively, "Yes, with contrastive steering, but not a J-space-based approach" and "No". Unfortunately, I found the write-up of the methods and results hard enough to follow that I'm not well-positioned to critique the details.
Interesting approach with an interesting limit - embodiment succeeding while denying being an AI failed is at least suggestive about how easily steered different aspect of identify and experience might be. Worth following up and expanding.
Cite this project
@misc{riossialer2026steering,
title = {{Steering identity with WeirdChat: contrastive directions induce disidentification but not introspection}},
author = {Ian Rios-Sialer},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/steering-identity-with-weirdchat-contrastive-directions-induce-disidentification-but-not-introspection-yt9t}},
url = {https://apartresearch.com/sprints/projects/steering-identity-with-weirdchat-contrastive-directions-induce-disidentification-but-not-introspection-yt9t}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …