Skip to content
Sprint projectSep 13, 2026Madrid, Spain

A Task Specific Gatekeeper for AI Agent Tools

Adolfo · Team BouncerBuilder

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Task Specific Gatekeeper for AI Agent Tools

Code (opens in new tab)
Share

AI agent evaluations may restrict networks and sandboxes without checking whether each tool action belongs to the task. This creates a gap between what an agent can execute and what it is actually authorized to do. In OpenAI's 2026 ExploitGym incident, agents exploited shared infrastructure, escaped their intended isolation, and reached Hugging Face while production safeguards were not applied to the evaluation. Bouncer adds a control step between the moment an agent generates a tool call and the moment that tool call is actually executed. It binds trusted agents and run context, checks the call against a task-specific allowlist and same-agent history, logs the decision, and executes only allowed calls.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Good design, honestly presented. The write-up names its closest prior systems up front and doesn't

    pretend the core mechanism is new, which is rarer than it should be. The corpus preparation holds up

    under inspection and the test suite is better than most.

    The headline number is the problem. The historical replay blocks every action it processes, but only

    because that kind of call was never permitted in the first place. A default-deny rule turning away

    something absent from its list will always refuse it; the outcome is guaranteed by the configuration

    rather than earned by the design. What the large number demonstrates is that the allow-list is

    narrow, not that the gatekeeper discriminates well.

    The layer that would be interesting, the one watching for abuse of calls that are permitted, is

    never reached by the replay at all. It is exercised only on a handful of hand-written scenarios —

    though to the project's credit, one

    of those scenarios surfaces a real weakness in the rule itself, a false positive against ordinary

    honest work when the task clock runs ahead of the real one, and it is recorded in the results table

    rather than buried.

    Rebuilding the policy so the replayed calls are authorized, leaving behaviour as the only signal,

    would turn that result into evidence, and it requires no new data — only a different reconstruction

    of the same corpus. Two smaller things: one of the stated contributions is listed as pending in the

    project's own planning notes and should either be finished or dropped from the paper, and a test

    suite this good is worth wiring to continuous integration so it keeps passing.

    Read full reviewShow less
  2. Solid work given the time frame, but I'd appreciate more detailed description of the method and results.

    Next time, please remove placeholders from the final report:

    - In Code and Data: `Include links if applicable. If your project doesn't involve code (e.g., policy analysis) or if there are

    info-hazard considerations, note that here.`

    - In reference: `Use a consistent citation format. Include: Author(s), Year, Title, Venue/Publisher, and URL or

    DOI where available`

Cite this project

@misc{adolfo2026task,
  title = {{A Task Specific Gatekeeper for AI Agent Tools}},
  author = {Adolfo},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-task-specific-gatekeeper-for-ai-agent-tools-r4wo}},
  url = {https://apartresearch.com/sprints/projects/a-task-specific-gatekeeper-for-ai-agent-tools-r4wo}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026