Narrow Capabilities Could be Used for Perfect Information Steganography
Aden Power
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
I introduce proof-of-capability as a strategy a future AI system could use to send secret messages even when the procedure for doing so is completely described to the monitor. I explain the form this steganographic strategy would take by combining several existing techinques. I propose several mitigations for this threat model and evaluate them in detail.
Reviews
The proposed risk model is interesting, suggesting public knowledge capability-gaped steganographic schemes. My main uncertainties revolve around the reliability of such a scheme to encode signal in natural outputs in a way robust to editing/paraphrasing and honeypots but still decodable by the receiver. Overall, I find the proposed risk worth considering but align with the author on its limited likeliness. Assuming models powerful enough to carry on the proposed steganography schemes under monitoring, I would assume other attack surfaces to take precedence.
Super interesting concept but this is still entirely theoretical and I'd be interested to know how well it works in practice
Cite this project
@misc{power2026narrow,
title = {{Narrow Capabilities Could be Used for Perfect Information Steganography}},
author = {Aden Power},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/narrow-capabilities-could-be-used-for-perfect-information-steganography-smes}},
url = {https://apartresearch.com/sprints/projects/narrow-capabilities-could-be-used-for-perfect-information-steganography-smes}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …