Trusted Model Supervisor
Kalpesh Panchal · Team Trusted Model Supervisor
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Trusted Model Supervisor is a defensive AI control layer designed to monitor, audit, and flag potentially harmful behavior from untrusted large language models. As increasingly capable AI systems are deployed in environments where misuse or misalignment can cause significant harm, operators currently lack lightweight, practical tools that provide real-time oversight. This project introduces an open-source, containerized monitoring pipeline consisting of a sandboxed untrusted model, a probe manager for query routing and telemetry collection, a trusted filtering layer for rule-based and classifier-emulated severity scoring, and a frontend dashboard enabling operators to observe incidents, replay interactions, and evaluate model behavior. The system logs all interaction data into a Postgres-backed audit trail, enabling reproducibility and post-hoc analysis. The MVP demonstrates that defensive control layers can be built efficiently using modern, lightweight tooling and containerized isolation, making it feasible to deploy safety supervision even in resource-constrained or rapidly changing environments.
Reviews
No public critique yet.
Cite this project
@misc{panchal2025trusted,
title = {{Trusted Model Supervisor}},
author = {Kalpesh Panchal},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/trusted-model-supervisor-8o2r}},
url = {https://apartresearch.com/sprints/projects/trusted-model-supervisor-8o2r}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …