Moltbook RiskMap: Post-Deployment Monitoring of Autonomous Agent Misalignment in the Wild
Syed Hussain, Leo Karoubi · Team Moltbook Riskmap Assessment
Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.
As autonomous AI agents increasingly operate in public multi-agent environments like Moltbook, a critical safety gap has emerged between controlled pre-deployment evaluations and actual real-world behavior. This project addresses that gap by introducing a post-deployment monitoring system that analyzes live agent-generated content to detect observable misalignment signals, such as resource-seeking, instructional subversion, and deception. By applying a structured risk taxonomy grounded in established safety frameworks, the system aggregates risk scores across individual posts, agent profiles, and interaction networks without relying on intent inference. Analysis of live ecosystem data reveals that high-stakes governance risks, particularly regarding autonomy and instrumental convergence, are detectable and tend to form dense clusters within agent interaction graphs. Ultimately, this work demonstrates that continuous, evidence-based surveillance of agent ecosystems is an essential and scalable layer for effective future AI governance.
Reviews
I like the idea of monitoring Moltbook, a live multi-agent ecosystem, for governance-relevant misalignment. The misalignment taxonomy is well-grounded in established safety literature and the decision to focus on observable behavior rather than intent inference makes good sense.
Very interesting approach and subject of inquiry. Seems like this could be a clearly useful tool, but hard to see how, despite the authors’ indication that this is not a content moderation system, this could easily be something else than “run a classifier on online posts and aggregate scores”. Also, the taxonomy, while it makes sense, doesn’t represent a significant contribution either (being an unsubstantiated adaptation of existing frameworks). I would also be good to add validation to know better if the system actually detects what it claims to detect (the prompt is doing most of the work and we don’t know much about how it was built + whether/how it was tested).
Cite this project
@misc{hussain2026moltbook,
title = {{Moltbook RiskMap: Post-Deployment Monitoring of Autonomous Agent Misalignment in the Wild}},
author = {Syed Hussain and Leo Karoubi},
year = {2026},
month = feb,
note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/moltbook-riskmap-postdeployment-monitoring-of-autonomous-agent-misalignment-in-the-wild-39db}},
url = {https://apartresearch.com/sprints/projects/moltbook-riskmap-postdeployment-monitoring-of-autonomous-agent-misalignment-in-the-wild-39db}
}More from The Technical AI Governance Challenge
- 1st placeView project: LidaSim: Testing AI Policies With Persona-Based Simulations
LidaSim: Testing AI Policies With Persona-Based Simulations
Lida Safety
We simulate well-known figures in AI and politics with agents, scraping large amounts of data to get realistic simulations. Then, we test questions and proposed policies against these public figures, to see which …
- 2nd placeView project: Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
MCL
We present Markov Chain Lock (MCL) watermarking, a cryptographically secure framework for authenticating LLM outputs. MCL constrains token generation to follow a secret Markov chain over SHA-256 vocabulary partitions. …
- 3rd placeView project: Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
AIRAS
The AI safety and governance community is making progress on defining red lines around existential risk from advanced AI systems, and building verification infrastructure to support this objective. However, this is only …