Skip to content
Sprint projectSep 14, 2026West Lafayette

MARGIN Arena: Agents Under Resource Limits

Seunghyun Yoo · Team SYoo

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: MARGIN Arena: Agents Under Resource Limits

Code (opens in new tab)
Share

MARGIN Arena is an installable research environment for studying how AI agents behave under explicit resource constraints, including generated text, elapsed time, verification attempts, actions, and fresh starts. It connects existing tasks with compatible open models, enforces resource limits, and records agent decisions, outcomes, and optional internal measurements to support analysis of behaviors such as shortcutting, reporting errors, and rule violations.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper introduces an innovative and practical infrastructure tool by formalizing an AI agent "flight recorder" across five granular resource constraints (generated text, elapsed time, verification attempts, actions, and fresh starts). Studying how resource scarcity alters an agent's propensity to bypass rules or engage in reward hacking is highly relevant to real-world deployment safety. The environment's integration with popular mechanistic interpretability libraries (such as TransformerLens and SAE Lens) provides a useful foundation for deeper safety forensics.

    The main limitation is that the paper serves primarily as an environment announcement rather than a completed empirical study. The experimental evaluation is extremely minimal, consisting of only a four-run illustrative example using Qwen and QuixBugs.

    Furthermore, as documented in Appendix C, the framework's internal measurement feature suffered a significant technical setback: all replay comparisons failed the preset numerical agreement threshold, meaning the captured internal states cannot yet be reliably used to interpret original agent decisions. Future versions must resolve these activation synchronization issues and validate the arena across a wider suite of automated hacking benchmarks.

    Read full reviewShow less
  2. A real tool rather than a sketch of one. Independent resource limits with proper reserve-and-settle accounting, spend that survives restarts, a tamper-evident journal, sandbox isolation, working continuous integration, a published package. It wraps existing evaluation tasks instead of reinventing them, and refuses to invent its own success labels — the kind of restraint that makes a tool safe for other people to build on. The validation checks the paper claims are real and present in the repository.

    Nothing empirical ships with it, though. The model runs described in the appendices, and the interpretability work alongside them, have no artifacts in the repository at all; the run directory is excluded from version control. Everything a reader can verify is infrastructure, and nothing shows the environment doing its job against an actual model. Committing a single complete run as a fixture — journal, config, outputs, resource ledger — would close that, and it is a small change relative to what has already been built. The journal is the feature a reader would most want to see working, so having no instance of one to inspect is the gap most worth filling.

    Read full reviewShow less

Cite this project

@misc{yoo2026margin,
  title = {{MARGIN Arena: Agents Under Resource Limits}},
  author = {Seunghyun Yoo},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/margin-arena-agents-under-resource-limits-630i}},
  url = {https://apartresearch.com/sprints/projects/margin-arena-agents-under-resource-limits-630i}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026