TDMF: Trajectory Deviation Monitor with Forensic Deep-Dive
Abubakar Abdulfatah · Team Rakabuba
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
TDMF is a trajectory-aware AI safety monitor that detects coordinated multi-step attacks — sequences where each individual action is locally innocent, but the full trajectory constitutes a staged exploit (reconnaissance → payload → exfiltration).
Current action-only monitors score each step in isolation. TDMF maintains a rolling forensic audit ledger and scores actions against a constraint envelope (what the agent is categorically forbidden from doing), not a prediction of what it will do. When violations accumulate, a forensic deep-dive reconstructs causal chains across the full history.
Reviews
This is a cool idea and points out a genuine flaw with common monitors that aren't given context on future actions. These common monitors are shown all past actions up to the action so they aren't entirely stateless, and in LinuxArena, you need these action level monitors to handle live failure side tasks. But for the existing literature, this author points out a good point that having a monitor not see into the future when evaluating an action is missing alpha. Also they propose hierarchical monitoring which has been explored in the literature, but cool to see anyway.
AI was used a bit to much to write the code and paper. It made the paper too long, with too much detail, and if you look at the codebase, the transcripts used to evaluate the monitor were hardcoded and very minimal: https://github.com/abubakar-xyz/tdmf-ai-control/blob/main/tdmf/eval/transcripts.py . Would have been cool to have seen this done with real agent transcripts (e.g. those available on linuxarena.ai).
Read full reviewShow less
Nice implementation of a multi-turn policy, and nice approach.
One minor engineering nit: this is poorly documented, but Luthien policy's are (supposed to be) ~stateless, as the same policy can be acting on many requests at once and you usually don't want those results interfering with each other - there are ways to persist state across turns that are session-specific, but they're poorly documented. This is mostly a criticism of Luthien's documentation at the moment, rather than your implementation.
Cite this project
@misc{abdulfatah2026tdmf,
title = {{TDMF: Trajectory Deviation Monitor with Forensic Deep-Dive}},
author = {Abubakar Abdulfatah},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/tdmf-trajectory-deviation-monitor-with-forensic-deepdive-rmql}},
url = {https://apartresearch.com/sprints/projects/tdmf-trajectory-deviation-monitor-with-forensic-deepdive-rmql}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …