Point of No Return: Does Concept Injection Break Reasoning Models?

Akshata Bhat

We test whether concept injection, the technique used to probe LLM introspection, remains safe when applied to reasoning models generating extended chain-of-thought, rather than the short single-turn outputs it was validated on. Across 7 open-weight reasoning models, 4 concept directions, a random-noise control, and 4 injection depths, we find that an injection strength validated as safe for short-form concepts collapses valid completion to 0% in 105 of 112 conditions. A token-level removal ablation reveals a model-dependent "point of no return," varying by more than an order of magnitude across models, past which removing the injection no longer restores a valid completion (in one model the output is byte-for-byte identical whether the perturbation continues or stops). We show this failure is largely generic to perturbation magnitude rather than concept-specific, present an initial mechanistic account (entropy collapse) on two models, and argue it is both a hidden confound for concept-injection introspection benchmarks and a warning for real-time safety monitors that assume removing a disturbance's source stops its effects.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is an ambitious and technically strong sprint project that identifies an important introspection-evaluation confound: concept injection can prevent reasoning models from producing valid completions, and removing the injection may not restore the trajectory. The breadth across seven models, four concepts, multiple layers, random-vector controls, removal experiments, and entropy analysis makes the work valuable for benchmark designers. However, the “short-form-safe” characterization should be narrowed because the operating point was selected using Qwen3-8B and produced poor or zero short-form completion in several other models. The point-of-no-return estimates also rely on one prompt and deterministic trajectories, while a single random direction is insufficient to characterize generic perturbations. Future work should calibrate strength separately for each model, test multiple random directions and prompts, add sampled repetitions, and distinguish persistent effects of generated context from changes in internal state. Overall, this is an original and highly useful extension of concept-injection research,

Summary and research question

This project asks whether concept injection, previously used in short-form introspection experiments, remains stable during extended reasoning, and whether a reasoning model can recover once the intervention is removed.

Across seven open-weight reasoning models, four concept directions, several injection depths, and removal-time ablations, the authors find a striking and highly reproducible pattern: sustained concept injection often causes extended generations to collapse or fail to complete. They further show that after sufficient exposure, simply stopping the intervention may not restore a valid trajectory, and provide an initial entropy-based explanation for this persistence.

Strengths

- Interesting and practically relevant benchmark-design question: introspection experiments should distinguish failure to detect an injected concept from failure to generate a valid response at all.

- Strong cross-model breadth, with seven reasoning models spanning several architecture families.

- Good experimental breadth across multiple concept directions, random-direction controls, layer depths, and removal times.

- The transcript examples provide convincing qualitative evidence that the observed failures are genuine generation degeneration rather than merely formatting artifacts.

- The removal experiment is a creative first attempt to study whether perturbation effects can become self-sustaining during autoregressive reasoning.

- The entropy analysis is a promising initial mechanistic lead, and the authors appropriately acknowledge that it is only clearly established on one model.

- The recommendation to report completion failure separately from detection rates is immediately useful for future introspection benchmarks.

Limitations

-The paper describes the operating point as short-form-safe, but this is only strongly supported for some models; several already show substantial short-form degradation at the same strength.

-Turning the steering hook off does not necessarily erase its earlier effects from the generated prefix or cached model state, so the removal experiment is best interpreted as evidence of persistent trajectory effects rather than a fully isolated internal “point of no return.”

- The model-specific recovery boundaries are currently based on very few deterministic trajectories, so they should be treated as preliminary windows rather than precise thresholds.

- Random vectors are nearly as destructive in several models, suggesting that part of the phenomenon may reflect large residual-stream perturbations in general, not concept semantics specifically.

- The strength sweep is exploratory and sparse at several values, so the boundary between informative steering and destabilization remains under-characterized.

- The broader implications for prompt injection, monitoring, and tool-output sanitization are interesting hypotheses, but require more direct experiments before they can be generalized confidently.

Overall assessment

This is a creative and worthwhile sprint project with a clear empirical signal. The strongest result is that activation injections that appear usable in short-form settings can interact very differently with extended reasoning, and that introspection benchmarks need to measure generation stability explicitly.

The “point of no return” framing is an intriguing hypothesis rather than a fully established mechanism at this stage. A particularly strong follow-up would validate safe injection strengths separately for each model, replicate removal curves over multiple prompts and concepts, and reconstruct the unsteered state after removal where possible. That would turn an interesting robustness phenomenon into a much cleaner causal result.

Cite this work

@misc {

title={

(HckPrj) Point of No Return: Does Concept Injection Break Reasoning Models?

},

author={

Akshata Bhat

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923