Hidden in Plain Sight: Representational Adversarial Steganography in Colluding LLMs
Vainateya Rangaraju, Srujana Medicherla, Denis Lim, Guillaume Zahnd, Anna Konovalenko, Igor Pereverzev · Team RASCAL
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
We investigate whether Large Language Models (LLMs) can develop representational steganography: the ability to encode and transmit hidden signals within internal representations that are not readily detectable by oversight mechanisms. We apply this to the threat models of deception and collusion in multi-agent settings. We show that this capability emerges under adversarial pressure, and discuss the robustness of our experiments, and what they might show for existing machine oversight regimes.
Reviews
Interesting work on representational steganography for multi-agent control. But I have two main concerns:
1. The threat model: it's not clear why two agents in the wild would share representations through an explicit latent channel; real multi-agent communication happens via text/APIs, so the setup feels contrived
2. The evaluation - the key claim is that communication survives adversarial pressure from the monitor, but the available results mostly show Experiment 1 (communication without evasion), and ROC/safety curves would have been much more interpretable than point accuracies for telling the safety story about monitor evasion tradeoffs.
A well-presented, technical implementation and exploration of steganography in a shared-activation space situation. The rank-8 constraint is a smart bottleneck. I thought the figures were very useful.
The motivation of this felt like the notable weak point. Current multi-agent systems communicate through text and not through shared residual streams, so the practical relevance of this threat model isn't clear to me. This ultimately detracted from the underlying idea and the work - which I think were both quite good. Also, this result required deliberate trial-and-error training pressure, and did not arise spontaneously, which makes this a less pressing safety concern.
I still think there are fruitful areas to explore here, if you wanted to take this further. There may be an argument to be made about multi-headed or mixture-of-experts architectures where components do share internal representations, and whether selection pressures in those settings could produce something analogous to your training dynamics. I'd be interested in your take on that!
Read full reviewShow less
Cite this project
@misc{rangaraju2026hidden,
title = {{Hidden in Plain Sight: Representational Adversarial Steganography in Colluding LLMs}},
author = {Vainateya Rangaraju and Srujana Medicherla and Denis Lim and Guillaume Zahnd and Anna Konovalenko and Igor Pereverzev},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/hidden-in-plain-sight-representational-adversarial-steganography-in-colluding-llms-sekh}},
url = {https://apartresearch.com/sprints/projects/hidden-in-plain-sight-representational-adversarial-steganography-in-colluding-llms-sekh}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …