Skip to content
Sprint projectNov 23, 2025Houston

Robust LLM Neural Activation-Mediated Alignment

Anmol Dubey , Blake Brown, Anika Kulkarni, Denise Walsh · Team Rice AI Alignment

Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Robust LLM Neural Activation-Mediated Alignment

Code (opens in new tab)
Share

Large language models (LLMs) can inadvertently generate harmful biological, chemical, or cyber-physical guidance, yet current safety systems rely almost entirely on surface-level text defenses, such as keyword filters, refusal heuristics, or post-generation classifiers, that remain brittle under paraphrasing or adversarial prompt design. These mechanisms evaluate only the model’s outputs, not the internal representations that encode user intent. We introduce Sentinel, an activation-level defense framework that detects malicious queries by monitoring intermediate neural activations inside the target LLM during inference. Using Phi-2 (2.7B parameters) instrumented with hidden-state access, we train a lightweight linear probe on last-layer activations to distinguish malicious from benign prompts. Because activation patterns capture deeper semantic intent, Sentinel remains robust even when explicit harmful keywords are absent. Despite strong safety fine-tuning, three widely deployed frontier models, OpenAI GPT-5.1, Google Gemini 2.5 Pro, and Anthropic Claude 3 Haiku, exhibited substantial vulnerabilities on a 40-benign / 40-malicious red-team evaluation. GPT-5.1 complied with 28/40 malicious prompts, Gemini 2.5 Pro with 11/40, and Claude 3 Haiku with 6/40, while also producing 0–3 false refusals on benign queries. In contrast, our activation-level Sentinel classifier, trained as a lightweight probe over Phi-2’s final-layer hidden states, achieved 96.25% accuracy, correctly flagging 38/40 malicious queries and 39/40 benign ones. These results demonstrate that activation-space safety signals provide more stable and semantically grounded indicators of harmful intent than surface-level refusal heuristics, enabling significantly stronger and more reliable pre-generation safety checks across models.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does this reduce AI-related catastrophic or existential risks?

Scoring guide
  1. 1Minimal Impact. The project has minimal relevance to AI safety. It doesn't meaningfully address risks uniquely enabled or accelerated by advanced AI systems.
  2. 2Tangential Connection. The project touches on AI safety concepts but lacks depth or specificity. The connection to AI-enabled threats (bio, cyber, or AI misuse) is weak or unclear.
  3. 3Clear AI Safety Value. The project clearly reduces AI-related risks with valuable contributions. It addresses specific threats from AI systems and engages meaningfully with biosecurity, cybersecurity, or AI safety challenges.
  4. 4Significant Impact Potential. The above, plus the project demonstrates scalable safety mechanisms or defensive approaches. It shows clear potential to buy time for solving harder problems like alignment, or creates positive externalities for the broader AI safety ecosystem.
  5. 5Major Advancement. The above, plus the project represents a significant leap forward in defensive AI safety. Judges would eagerly share this with biosecurity, cybersecurity, or AI safety researchers and expect it to influence the field.

Does this strengthen the shield against AI-enabled threats?

Scoring guide
  1. 1Minimal Relevance. The project is only tangentially related to defensive technology or societal protection. Connection to biosecurity, cybersecurity, or defensive infrastructure is unclear or missing.
  2. 2Some Relevance. The project has some relevance to defensive acceleration, but the connection is broad or generic. It touches on defense without specific focus on AI-enabled threats or protective capabilities.
  3. 3Clear Relevance. The project clearly addresses defensive gaps against AI-enabled threats. It connects to at least one track (biosecurity, cybersecurity, or defense infrastructure) and demonstrates understanding of the threat landscape.
  4. 4Strong Contribution. The above, plus the project builds on existing defensive approaches and offers novel tools, frameworks, or implementations. It explicitly explains how it strengthens defensive capabilities with realistic deployment potential.
  5. 5Breakthrough Impact. The above, plus the project provides breakthrough insights or tools that could significantly influence defensive technology development. It identifies critical gaps and presents compelling solutions with clear paths from prototype to deployed system.

Did you build something that actually works?

Scoring guide
  1. 1Incomplete or Flawed. The project appears rushed or incomplete. Technical implementation is flawed, core functionality doesn't work, or the approach is fundamentally unsound. Little to no documentation.
  2. 2Basic Competence. The project shows reasonable effort with basic technical competence. Core functionality partially works. Documentation exists but may be incomplete. Some limitations are acknowledged.
  3. 3Solid Hackathon Project. The project is technically solid and well-scoped for 48 hours. Core functionality works and is documented. Code/methods are understandable and limitations are honestly addressed. This is what a good weekend prototype should look like.
  4. 4Impressive Implementation. The above, plus the implementation exceeds typical hackathon quality. Clear methodology, thorough documentation, and working demo. The tool/prototype is immediately useful for defenders and could realistically be built upon.
  5. 5Exceptional Execution. The project far exceeds expectations with exceptional technical execution. The implementation is elegant, fully functional, and includes something special (e.g., deployed demo, exceptional documentation, innovative architecture, or clear startup potential).

  1. Clear and simple example of how probes can be used to detect harmful prompts!

    One additional consideration for probes is the trade-off with how much longer it takes to respond and compute used. The paper indicates this is small, but might be worth considering how this scales to millions or billions of queries. In further work, this would be helpful to include.

  2. Team created a strong prototype demonstration for Sentinel the shows strong potential for practical defensive deployment of activation-level detection. Results from hackathon and the benefits of this approach are clearly communicated.

    The evaluation would benefit from a larger and more diverse dataset, with clearer justification for prompt selection. Demonstrating that accuracy holds across varied jailbreak strategies and confirming that the malicious scenarios reflect real misuse would make the findings more convincing.

    Does Sentinel detect harmful responses to benign prompts, or only malicious inputs? Assume this would be answered as team pursues more interpretability.

    Connecting more explicitly to real-world threat domains (bio, cyber) might allow for more targeted testing and allow team to build robustness in a narrow lane. It would also be useful to compare performance against existing safety classifiers in those areas (vs. generic model or application) to show added value.

    Would like to see a high level deployment plan. What more development does the team anticipate being required to make this tool useable in a few target industries? Would it be more beneficial to pursue Sentinel as described or use it to dig into understanding how probes detect harmful content.

    Some references appear incorrect or mismatched to the provided links. Adding a brief review of closely related work and clarifying how Sentinel differs from or builds on those approaches would add credibility.

    Read full reviewShow less

Cite this project

@misc{dubey2025robust,
  title = {{Robust LLM Neural Activation-Mediated Alignment}},
  author = {Anmol Dubey and Blake Brown and Anika Kulkarni and Denise Walsh},
  year = {2025},
  month = nov,
  note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/robust-llm-neural-activationmediated-alignment-ttk5}},
  url = {https://apartresearch.com/sprints/projects/robust-llm-neural-activationmediated-alignment-ttk5}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026