The State Is Not Enough

Krishna Sandoval Cambranis

Safety auditing tools for language models, such as backdoor detection and causal tracing, were built almost entirely for the Transformer and its attention mechanism. Selective state-space models like Mamba are now deployed at scale but carry information through a recurrent state instead of attention, so it is unclear whether these tools transfer. We study Mamba-790M and deliver two results. First, we show the selective time-step parameter that governs the model's memory is interpretable: it assigns write strength by token type and is decoupled from token frequency. Second, we plant a backdoor inside the recurrent state and causally localize it by ablating the full state pathway. We uncover a saturated residual that survives even exhaustive ablation, proving the backdoor travels in part by a second route that bypasses the state entirely. This exposes a concrete blind spot: defenses that target only the recurrent state cannot fully neutralize a planted backdoor.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This paper earns real credit for doing something technically difficult: adapting interpretability interventions to a non-Transformer architecture and applying them to a safety-relevant behavior rather than a descriptive analysis of factual recall. The core finding that exhaustively ablating the trigger's entire recurrent-state pathway suppresses most but not all of a planted backdoor, with the residual plateauing under band widening is a genuine result with direct implications for anyone building safety tooling for SSMs or SSM-Transformer hybrids.

The observational finding on ∆ (content-dependent memory write policy, decoupled from token frequency) is a clean contribution and a necessary prerequisite for the causal work. The reproducibility detail about needing to uninstall the CUDA kernels to expose ∆ through hooks is the kind of hard-won practical knowledge that's easy to miss and valuable to document. The main statistical concern is one the author states directly and repeatedly: everything is a point estimate from a single seed, a single 790M model, and a single trigger-target pair. The 0.28 residual and 42% single-pathway neutralization should be read as illustrative rather than precise. That's an honest framing, but it also means the central claim that a non-state pathway carries part of the backdoor rests on a single experiment. A second seed and a second trigger pair would substantially strengthen confidence. The pre-registered prediction about content-sensitivity that ran in the opposite direction is worth a sentence or two of more explicit interpretation rather than just a notation in "what did not work." That's a finding too, not just a failure. Strong work overall given the scope constraints.

It was hard to identify how valuable the findings of the article were. Being time-constrained in reviewing, I could not follow through the full logic of the experiments, and I would have liked to have a clean comparison against existing methods and ideally some clean specification of the setup to study. The paper needed either a bigger effort to engage with existing methods (eg., a table comparing its approach against numbers from other papers) or to excel in writing, such that from the abstract it was clearer what the methods and setup were, or at least how they could be easily understood in relation to other papers.

The article was also really wordy, which made it hard to follow and assess. Overall, it was possibly a good contribution, but the lack of legibility + time constrain made it hard to assses that

I don't think the selected problem is relevant. Moreover, the key empirical question (how well do white-box classifiers work on the Mamba architecture) is not answered clearly.

Great contribution to interpretability in SSM, the findings are novel and timely. Also very well written, however the writing could benefit from more visual aids or in-depth explanation of key technical concepts While it demonstrates the vulnerability effectively, more diverse attacks would strengthen the execution by showing broader generalizability

Cite this work

@misc {

title={

(HckPrj) The State Is Not Enough

},

author={

Krishna Sandoval Cambranis

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.