Skip to content
Sprint projectJun 21, 2026Mexico, Yucatam

The State Is Not Enough

Krishna Sandoval Cambranis · Team Krishna Team

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Safety auditing tools for language models, such as backdoor detection and causal tracing, were built almost entirely for the Transformer and its attention mechanism. Selective state-space models like Mamba are now deployed at scale but carry information through a recurrent state instead of attention, so it is unclear whether these tools transfer. We study Mamba-790M and deliver two results. First, we show the selective time-step parameter that governs the model's memory is interpretable: it assigns write strength by token type and is decoupled from token frequency. Second, we plant a backdoor inside the recurrent state and causally localize it by ablating the full state pathway. We uncover a saturated residual that survives even exhaustive ablation, proving the backdoor travels in part by a second route that bypasses the state entirely. This exposes a concrete blind spot: defenses that target only the recurrent state cannot fully neutralize a planted backdoor.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper earns real credit for doing something technically difficult: adapting interpretability interventions to a non-Transformer architecture and applying them to a safety-relevant behavior rather than a descriptive analysis of factual recall. The core finding that exhaustively ablating the trigger's entire recurrent-state pathway suppresses most but not all of a planted backdoor, with the residual plateauing under band widening is a genuine result with direct implications for anyone building safety tooling for SSMs or SSM-Transformer hybrids.

    The observational finding on ∆ (content-dependent memory write policy, decoupled from token frequency) is a clean contribution and a necessary prerequisite for the causal work. The reproducibility detail about needing to uninstall the CUDA kernels to expose ∆ through hooks is the kind of hard-won practical knowledge that's easy to miss and valuable to document. The main statistical concern is one the author states directly and repeatedly: everything is a point estimate from a single seed, a single 790M model, and a single trigger-target pair. The 0.28 residual and 42% single-pathway neutralization should be read as illustrative rather than precise. That's an honest framing, but it also means the central claim that a non-state pathway carries part of the backdoor rests on a single experiment. A second seed and a second trigger pair would substantially strengthen confidence. The pre-registered prediction about content-sensitivity that ran in the opposite direction is worth a sentence or two of more explicit interpretation rather than just a notation in "what did not work." That's a finding too, not just a failure. Strong work overall given the scope constraints.

    Read full reviewShow less
  2. It was hard to identify how valuable the findings of the article were. Being time-constrained in reviewing, I could not follow through the full logic of the experiments, and I would have liked to have a clean comparison against existing methods and ideally some clean specification of the setup to study. The paper needed either a bigger effort to engage with existing methods (eg., a table comparing its approach against numbers from other papers) or to excel in writing, such that from the abstract it was clearer what the methods and setup were, or at least how they could be easily understood in relation to other papers.

    The article was also really wordy, which made it hard to follow and assess. Overall, it was possibly a good contribution, but the lack of legibility + time constrain made it hard to assses that

  3. I don't think the selected problem is relevant. Moreover, the key empirical question (how well do white-box classifiers work on the Mamba architecture) is not answered clearly.

  4. Great contribution to interpretability in SSM, the findings are novel and timely. Also very well written, however the writing could benefit from more visual aids or in-depth explanation of key technical concepts While it demonstrates the vulnerability effectively, more diverse attacks would strengthen the execution by showing broader generalizability

Cite this project

@misc{cambranis2026state,
  title = {{The State Is Not Enough}},
  author = {Krishna Sandoval Cambranis},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-state-is-not-enough-26ax}},
  url = {https://apartresearch.com/sprints/projects/the-state-is-not-enough-26ax}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026