Point of No Return: Does Concept Injection Break Reasoning Models?
Akshata Bhat
We test whether concept injection, the technique used to probe LLM introspection, remains safe when applied to reasoning models generating extended chain-of-thought, rather than the short single-turn outputs it was validated on. Across 7 open-weight reasoning models, 4 concept directions, a random-noise control, and 4 injection depths, we find that an injection strength validated as safe for short-form concepts collapses valid completion to 0% in 105 of 112 conditions. A token-level removal ablation reveals a model-dependent "point of no return," varying by more than an order of magnitude across models, past which removing the injection no longer restores a valid completion (in one model the output is byte-for-byte identical whether the perturbation continues or stops). We show this failure is largely generic to perturbation magnitude rather than concept-specific, present an initial mechanistic account (entropy collapse) on two models, and argue it is both a hidden confound for concept-injection introspection benchmarks and a warning for real-time safety monitors that assume removing a disturbance's source stops its effects.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Point of No Return: Does Concept Injection Break Reasoning Models?
},
author={
Akshata Bhat
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


