Panopticon
Sanchayan Ghosh · Team Panopticon
Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.
A proof of chain verifier for AI Use, determining if the AI has been tampered with and if so, at what stage.
Also detect anomalous outputs in the LLM, and flag them to the user, organizing the LLM into dangerous, confusing or safe inputs.
Analyze the LLM's activation states to determine LLM's confusion on seeing malicious prompts helping users craft more malicious prompts to analyze the LLM
Reviews
The project tries to address hardware provenance, activation-based safety monitoring, and automated red-teaming simultaneously. All of these are important topics but each of them is a major research area with substantial existing literature.
My main suggestion would be to pick one problem and go deep. Try to understand what is known about the problem and where the real gaps lie. For example, for the activation monitoring direction (Layer 2) there exists a lot of highly relevant work on representation engineering and linear probes for safety-relevant features.
The next most important thing is to take evaluation much more seriously: Evaluate whether your method achieves reasonable results and really stress-test your findings. A convincing evaluation of one component is better than a full-stack demo with lacking validation.
The team built an impressive prototype. The core problem is that the system is not evaluated in any real way. My recommendation would be to pick the most promising component, run it against a concrete set of prompts, and show it catches something a simpler baseline doesn't.
Cite this project
@misc{ghosh2026panopticon,
title = {{Panopticon}},
author = {Sanchayan Ghosh},
year = {2026},
month = feb,
note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/panopticon-ja2g}},
url = {https://apartresearch.com/sprints/projects/panopticon-ja2g}
}More from The Technical AI Governance Challenge
- 1st placeView project: LidaSim: Testing AI Policies With Persona-Based Simulations
LidaSim: Testing AI Policies With Persona-Based Simulations
Lida Safety
We simulate well-known figures in AI and politics with agents, scraping large amounts of data to get realistic simulations. Then, we test questions and proposed policies against these public figures, to see which …
- 2nd placeView project: Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
MCL
We present Markov Chain Lock (MCL) watermarking, a cryptographically secure framework for authenticating LLM outputs. MCL constrains token generation to follow a secret Markov chain over SHA-256 vocabulary partitions. …
- 3rd placeView project: Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
AIRAS
The AI safety and governance community is making progress on defining red lines around existential risk from advanced AI systems, and building verification infrastructure to support this objective. However, this is only …