Skip to content
Sprint projectMar 22, 2026Chennai, India

Layered Scope-Boundary Monitoring: Cost-Aware Control for Untrusted AI Agents

Rahul Kumar · Team DeoGaze

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Layered Scope-Boundary Monitoring: Cost-Aware Control for Untrusted AI Agents

Recording (opens in new tab)Code (opens in new tab)
Share

This project builds and evaluates a hierarchical control protocol for untrusted AI agents, designed to intercept covert sabotage attempts while preserving the agent's ability to complete legitimate work. The core research question is whether monitoring cost should be uniform across all agent actions or matched to the type of threat each action poses. The answer, validated empirically across two independent control settings, is that attack surface type determines which monitoring layer is necessary. The protocol routes each agent action through layers of increasing cost. A zero-cost deterministic scope-boundary filter checks whether the action targets system resources outside what the task legitimately requires. This layer is task-adaptive, meaning it derives scope boundaries automatically from the task description at runtime rather than using hardcoded rules, which is why it achieves zero false positives across 369 monitored honest actions in two different settings. If the deterministic layer does not block, a cheap LLM sentinel scores the action's suspicion level, and ambiguous cases escalate to a stronger model. Additional features include honeypot credential injection for exfiltration detection, latency equalization to prevent timing side-channels, and a trajectory buffer that periodically checks whether individually clean actions are combining into a coordinated attack pattern. The key empirical finding is that the deterministic layer alone catches 4.7% of sophisticated human-crafted attacks at zero LLM cost, while adding the LLM layer raises detection to 90.6% on 106 adversarial trajectories from LinuxArena's GreatStraj dataset. The marginal contribution of the deterministic layer on code-level sabotage is only one percentage point over LLM monitoring alone, which is expected because code-level attacks do not touch system resources. The architectural value of the deterministic layer is cost routing on system-level threats, not detection improvement on harder attacks. On BashArena, which features system-level side tasks including privilege escalation and firewall disabling, the full protocol achieves 54% higher usefulness than the defer-to-trusted baseline while maintaining 87.5% safety and a 0.34% false positive rate. The project also includes a calibrated comparison between blind and informed monitoring approaches. Using FP-rate-calibrated thresholds derived from honest trajectory runs rather than fixed thresholds, the blind hierarchical monitor outperforms the informed baseline at moderate and generous audit budgets but underperforms at strict budgets, with the crossover depending on how aggressively a deployment is willing to flag suspicious trajectories for review. All results are fully reproducible. The repository includes the complete protocol implementation, evaluation scripts, pre-extracted scores, and a single verification script that confirms every number in the paper. Raw evaluation data is available separately. Full reproduction costs approximately fifteen dollars in API credits. The work was conducted as part of the AI Control Hackathon 2026 and received feedback from the Redwood Research team, which informed several methodological improvements including the switch from fixed-threshold to FP-rate-calibrated monitor comparisons and an honest reframing of the deterministic layer's contribution. The findings are being extended toward a full research paper with additional controlled experiments on the blind versus informed monitoring question and model scaling behavior.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper implements a multi-level approach to monitoring untrusted agents. The first layer (L0) is algorithmic, while L1 and L2 are LLM-based (L2 is stronger). The goal is to reduce cost by only relying on strong (expensive) monitors when cheap ones fail to detect it.

    I think that's an interesting approach to explore, but LinuxArena is not the right setting for that. In LinuxArena capable agents who know they are monitored are evaluated, and they know they only have one shot. In this setting most agent actions are innocuous, with very few well-crafted attacks.

    The approach implemented in this paper spends resources on benign actions by going through L0, L1, and L2 in every step. The paper would be much stronger if it explored layered monitoring the other way around: weaker monitors escalate to stronger ones in case of doubt (so benign actions are cheaper to evaluate rather than more expensive).

    The approach from this paper might be useful for settings that face a high frequency of low-quality attacks, and I'm curious to see how it fares there.

    Note: The paper doesn't seem to contain results from experiments including L2 (L0 and L1 only).

    Read full reviewShow less
  2. I like the idea of using more traditional security techniques and adding more layers to blue team protocols that aren't LLM based. I like the transperancy of your code/reasoning in your work.

    I'd probably use more samples for DTT. Overall, I think this is quite interesting work. The other place to improve (which is understandable given time contraints) is doing some iteration of red team to try and attack against this blue team protocol.

    I overall quite like this project.

Cite this project

@misc{kumar2026layered,
  title = {{Layered Scope-Boundary Monitoring: Cost-Aware Control for Untrusted AI Agents}},
  author = {Rahul Kumar},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/layered-scopeboundary-monitoring-costaware-control-for-untrusted-ai-agents-9dr7}},
  url = {https://apartresearch.com/sprints/projects/layered-scopeboundary-monitoring-costaware-control-for-untrusted-ai-agents-9dr7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026