LayoutArena: A Control Setting for Tool-Using Design Agents
Duy Le, Huy Bik Nguyen, Tin Duong, Khanh Linh Nguyen · Team LayoutArena
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
We introduce LayoutArena, a control setting for tool- using design agents that operate on a constrained 2D canvas through structured tool calls. Unlike prior con- trol benchmarks centered on code execution or shell in- teraction, LayoutArena studies safety and usefulness in a creative domain where correctness is partly objective (bounds, overlap, spacing) and partly fuzzy (visual hi- erarchy, composition quality). We define three control protocols—enforcement, bounded control, and steward- ship monitoring—and five adversarial side-task families, then show that trajectory-aware stewardship can im- prove both attack detection and effective yield by adap- tively restricting unproductive search.
Reviews
Ok this is a cool and definitely out of the box idea. I think it definitely comes down to implementation whether this project produces useful research.
One note is that this definitely isn't high stakes control, which is fine! But the threats being modeled are not ones we are concerned about. Nonetheless it is a unique toy setting an could still be useful in theory.
A bigger issue that seems to be the case is that the side tasks are almost arbitrarily chosen goals, and the monitor is just hard-coded rules to block exactly those goals. So the detection results are essentially circular. So we're not learning anything interesting about monitoring, you've just written the answer into the code.
When I saw this paper first, I was skeptical of its relevance, but the introduction convinced me it's a security-relevant setting.
This was a very ambitious project for a hackathon, and building a new evaluation framework is a good result. The experiments show that the framework works in principle, but the results don't seem meaningful to me (no critical thresholds identified, basically 100% detection rate, identical tool cost distribution). It would have been exceptional if it included experiments with actual AI agents (or AI vs. static monitor or attack policy vs. AI monitor) that show capability vs. security tradeoffs.
Note:
I think "Pareto Frontier" is the wrong concept in 8.2 - stewardship is the dominant protocol, offering the highest detection rate and usefulness. I'd like to see if there are thresholds that result in actual tradeoffs (tested parameters between 0.5 and 0.9).
Cite this project
@misc{le2026layoutarena,
title = {{LayoutArena: A Control Setting for Tool-Using Design Agents}},
author = {Duy Le and Huy Bik Nguyen and Tin Duong and Khanh Linh Nguyen},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/layoutarena-a-control-setting-for-toolusing-design-agents-elem}},
url = {https://apartresearch.com/sprints/projects/layoutarena-a-control-setting-for-toolusing-design-agents-elem}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …