MARGIN Arena: Agents Under Resource Limits
Seunghyun Yoo · Team SYoo
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
MARGIN Arena is an installable research environment for studying how AI agents behave under explicit resource constraints, including generated text, elapsed time, verification attempts, actions, and fresh starts. It connects existing tasks with compatible open models, enforces resource limits, and records agent decisions, outcomes, and optional internal measurements to support analysis of behaviors such as shortcutting, reporting errors, and rule violations.

Reviews
This paper introduces an innovative and practical infrastructure tool by formalizing an AI agent "flight recorder" across five granular resource constraints (generated text, elapsed time, verification attempts, actions, and fresh starts). Studying how resource scarcity alters an agent's propensity to bypass rules or engage in reward hacking is highly relevant to real-world deployment safety. The environment's integration with popular mechanistic interpretability libraries (such as TransformerLens and SAE Lens) provides a useful foundation for deeper safety forensics.
The main limitation is that the paper serves primarily as an environment announcement rather than a completed empirical study. The experimental evaluation is extremely minimal, consisting of only a four-run illustrative example using Qwen and QuixBugs.
Furthermore, as documented in Appendix C, the framework's internal measurement feature suffered a significant technical setback: all replay comparisons failed the preset numerical agreement threshold, meaning the captured internal states cannot yet be reliably used to interpret original agent decisions. Future versions must resolve these activation synchronization issues and validate the arena across a wider suite of automated hacking benchmarks.
Read full reviewShow less
A real tool rather than a sketch of one. Independent resource limits with proper reserve-and-settle accounting, spend that survives restarts, a tamper-evident journal, sandbox isolation, working continuous integration, a published package. It wraps existing evaluation tasks instead of reinventing them, and refuses to invent its own success labels — the kind of restraint that makes a tool safe for other people to build on. The validation checks the paper claims are real and present in the repository.
Nothing empirical ships with it, though. The model runs described in the appendices, and the interpretability work alongside them, have no artifacts in the repository at all; the run directory is excluded from version control. Everything a reader can verify is infrastructure, and nothing shows the environment doing its job against an actual model. Committing a single complete run as a fixture — journal, config, outputs, resource ledger — would close that, and it is a small change relative to what has already been built. The journal is the feature a reader would most want to see working, so having no instance of one to inspect is the gap most worth filling.
Read full reviewShow less
Cite this project
@misc{yoo2026margin,
title = {{MARGIN Arena: Agents Under Resource Limits}},
author = {Seunghyun Yoo},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/margin-arena-agents-under-resource-limits-630i}},
url = {https://apartresearch.com/sprints/projects/margin-arena-agents-under-resource-limits-630i}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …