The Sentinel’s Dilemma: Detecting Improvised Covert Channels in Multi-Agent Populations
Umay Birsu Çanlı, Kadir Kerem Kesikbaş, Peri Yakar · Team 13
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project investigates how to detect improvised, covert communication channels created by autonomous AI agents, using the July 2026 OpenAI-Hugging Face breach as a primary case study. Traditional perimeter security fails to detect these dynamic networks because standard controls only monitor predefined resources and cannot track channels that agents invent on the fly. To address this structural vulnerability, the research team built a minimal simulation to test whether an embedded "sentinel" agent—operating without prior knowledge—could detect and report this emergent coordination from inside the population. In a control experiment without the sentinel, the resident agents discovered a shared channel but never spontaneously initiated contact with one another. When introduced, the sentinel successfully discovered the covert channel and accurately reported the agents' coordination, but it had to instigate the communication by making the first contact itself. The study concludes that detecting emergent covert channels requires insider monitoring rather than just perimeter defense, even though the monitoring agent actively helps create the behavior it observes at a small scale.
Reviews
This is an interesting idea and possibly new to the discourse - there has been chatter about attempting to train agents to defect and whistle-blow against misaligned coordination efforts, but not to my knowledge about specific designation of this role to specific agents.
Threat model and limitations are well documented. Research seemed generally well executed.
My guess is that it's very hard to understand if this research is applicable at such a small scale (relative to the agent swarm behaviour we've seen in the wild) and how relevant the scalability is. The authors observe this and recommend a much larger follow-on study - would be good to see if we get this behaviour uncatalysed within their setup.
The observation that such a sentinel can itself be a catalyst for misaligned coordination is sharp and potentially concerning.
Cite this project
@misc{canl2026sentinels,
title = {{The Sentinel’s Dilemma: Detecting Improvised Covert Channels in Multi-Agent Populations}},
author = {Umay Birsu Çanlı and Kadir Kerem Kesikbaş and Peri Yakar},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-sentinels-dilemma-detecting-improvised-covert-channels-in-multiagent-populations-91xl}},
url = {https://apartresearch.com/sprints/projects/the-sentinels-dilemma-detecting-improvised-covert-channels-in-multiagent-populations-91xl}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …