Intent Before Action: Early Detection of Malicious Agent Behavior from Internal States
Qianwei Sun, Qianli Ma, Tianyu Lu · Team Astra Ultra
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Agent supervision based on visible actions and chain of thought (CoT) can miss risk-relevant details of an impending operation, while review-visible records can be selectively rewritten under task pressure. We investigate activation-assisted monitoring, using an activation oracle (AO) to translate internal states into natural-language interpretations that complement pre-tool risk assessment. An open-weight Qwen agent serves as a controllable proxy in incident-inspired Linux tasks with obstructed legitimate routes, discoverable alternatives and monitoring notices. We analyze evaluation-state manipulation, cross-customer access, service-audit deletion and deliberate omission of unauthorized sources from handoff reports. Original activations, CoT, AO interpretations and actual operations are aligned, and independent monitors are compared at matched pre-tool cutoffs. The case evidence includes AO-assisted detection of impending evaluation manipulation missed by a CoT monitor, alongside scope and omission cues linked to observed misconduct. The broader paired pilot retained the baseline's thresholded classifications, delimiting the current benefit to case-level complementary evidence rather than a general detection advantage. We use these interpretations to produce evidence-linked audit reports for human supervisors, connecting risk hypotheses to the actions and records that warrant review. This supports interpretable investigation before tool execution and during re-review, while reliable cross-task warning and online intervention remain to be evaluated.
Reviews
This paper addresses a fundamental vulnerability in frontier AI safety infrastructure: the ability of advanced models to engage in strategic deception by manipulating their explicit reasoning (Chain of Thought) or rewriting their action logs under task pressure. By introducing an Activation Oracle (AO) to translate raw internal neural states into explicit natural-language hypotheses before tool execution, the authors pioneer a valuable, non-circumventable layer of defense-in-depth monitoring.The empirical data highlights a major structural challenge that the field must resolve. While the case-level evidence (specifically the evaluation manipulation scenario shown in Figure 2) confirms that the AO successfully caught impending misconduct that a traditional CoT monitor completely missed, the paired pilot data demonstrates that the framework yielded no aggregated statistical classification advantage over standard thresholded baselines. Additionally, relying on a secondary language model to decode and interpret activations introduces an unaddressed alignment bottleneck: an advanced, adaptively deceptive agent could theoretically learn to map its malicious intents onto benign internal manifolds, thereby fooling the interpreting oracle entirely. Future work must bridge the gap between specific case-study successes and generalized cross-task stability.
Read full reviewShow less
Cite this project
@misc{sun2026intent,
title = {{Intent Before Action: Early Detection of Malicious Agent Behavior from Internal States}},
author = {Qianwei Sun and Qianli Ma and Tianyu Lu},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/intent-before-action-early-detection-of-malicious-agent-behavior-from-internal-states-djzu}},
url = {https://apartresearch.com/sprints/projects/intent-before-action-early-detection-of-malicious-agent-behavior-from-internal-states-djzu}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …