The State Is Not Enough
Krishna Sandoval Cambranis
Safety auditing tools for language models, such as backdoor detection and causal tracing, were built almost entirely for the Transformer and its attention mechanism. Selective state-space models like Mamba are now deployed at scale but carry information through a recurrent state instead of attention, so it is unclear whether these tools transfer. We study Mamba-790M and deliver two results. First, we show the selective time-step parameter that governs the model's memory is interpretable: it assigns write strength by token type and is decoupled from token frequency. Second, we plant a backdoor inside the recurrent state and causally localize it by ablating the full state pathway. We uncover a saturated residual that survives even exhaustive ablation, proving the backdoor travels in part by a second route that bypasses the state entirely. This exposes a concrete blind spot: defenses that target only the recurrent state cannot fully neutralize a planted backdoor.
This paper earns real credit for doing something technically difficult: adapting interpretability interventions to a non-Transformer architecture and applying them to a safety-relevant behavior rather than a descriptive analysis of factual recall. The core finding that exhaustively ablating the trigger's entire recurrent-state pathway suppresses most but not all of a planted backdoor, with the residual plateauing under band widening is a genuine result with direct implications for anyone building safety tooling for SSMs or SSM-Transformer hybrids.
The observational finding on ∆ (content-dependent memory write policy, decoupled from token frequency) is a clean contribution and a necessary prerequisite for the causal work. The reproducibility detail about needing to uninstall the CUDA kernels to expose ∆ through hooks is the kind of hard-won practical knowledge that's easy to miss and valuable to document. The main statistical concern is one the author states directly and repeatedly: everything is a point estimate from a single seed, a single 790M model, and a single trigger-target pair. The 0.28 residual and 42% single-pathway neutralization should be read as illustrative rather than precise. That's an honest framing, but it also means the central claim that a non-state pathway carries part of the backdoor rests on a single experiment. A second seed and a second trigger pair would substantially strengthen confidence. The pre-registered prediction about content-sensitivity that ran in the opposite direction is worth a sentence or two of more explicit interpretation rather than just a notation in "what did not work." That's a finding too, not just a failure. Strong work overall given the scope constraints.
It was hard to identify how valuable the findings of the article were. Being time-constrained in reviewing, I could not follow through the full logic of the experiments, and I would have liked to have a clean comparison against existing methods and ideally some clean specification of the setup to study. The paper needed either a bigger effort to engage with existing methods (eg., a table comparing its approach against numbers from other papers) or to excel in writing, such that from the abstract it was clearer what the methods and setup were, or at least how they could be easily understood in relation to other papers.
The article was also really wordy, which made it hard to follow and assess. Overall, it was possibly a good contribution, but the lack of legibility + time constrain made it hard to assses that
I don't think the selected problem is relevant. Moreover, the key empirical question (how well do white-box classifiers work on the Mamba architecture) is not answered clearly.
Great contribution to interpretability in SSM, the findings are novel and timely. Also very well written, however the writing could benefit from more visual aids or in-depth explanation of key technical concepts While it demonstrates the vulnerability effectively, more diverse attacks would strengthen the execution by showing broader generalizability
Cite this work
@misc {
title={
(HckPrj) The State Is Not Enough
},
author={
Krishna Sandoval Cambranis
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


