Mechanistic Watchdog
Luis Cosio · Team SL5
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear “concept vectors” that capture truthfulness and high-risk domains. Using datasets like Facts-True-False, TruthfulQA and WMDP-Bio, we calibrate a deception / misuse direction in the residual stream of mid-layers (e.g., Llama-3.1-8B, Qwen-2.5-3B). During generation, the watchdog projects each new token’s hidden state onto these vectors, smooths the scores, and halts the model if the trajectory crosses a learned threshold—interdicting deceptive or bio-risky cognition before it fully materializes in text.
Our experiments show that a single truthfulness vector trained on generic true/false facts generalizes out-of-distribution to TruthfulQA, cleanly separating truthful controls from misconceptions and factual lies. We also prototype a bio-defense profile that reliably detects when the model is “in a biological regime” and are iterating toward a contrastively trained safe-vs-misuse bio probe. Mechanistic Watchdog is intended as a building block for def/acc: an internal-state monitoring layer that defenders can place in front of powerful models to reduce the risk of AI-enabled deception and assist in catching early signs of bio- or cyber-misuse.
Reviews
Impressive amount to build as a single person in such a short time!
For future work, I'd love to see more exploration on stress testing this method with more prompts that intentionally try to jailbreak the models to see whether this manages to properly detect the flagged behaviour.
Report provides a clear path to deployment. The prototype’s results are encouraging across models and tasks, and the performance overhead seems acceptable. I appreciate the clear explanation of where this tool sits in the broader AI-safety ecosystem.
I would like to see more early thinking on robustness against adversarial adaptation, as motivated actors may learn to hide or route around these internal signals.
It would also be helpful to discuss risks from accidental trigger events in high-stakes settings. If the system halts or blocks a model at the wrong moment, could that itself cause harm?
Cite this project
@misc{cosio2025mechanistic,
title = {{Mechanistic Watchdog}},
author = {Luis Cosio},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/mechanistic-watchdog-klay}},
url = {https://apartresearch.com/sprints/projects/mechanistic-watchdog-klay}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: GUARDIAN: Guarded Universal Architecture for Defensive Interpretation And traNslation
GUARDIAN: Guarded Universal Architecture for Defensive Interpretation And traNslation
Guardian team
GUARDIAN is a multi-stage, LLM-driven system to automate the translation of C codebases to memory-safe Rust. GUARDIAN promotes defense acceleration at-scale by guiding an LLM transpiler with dependency graph …