A Task Specific Gatekeeper for AI Agent Tools
Adolfo · Team BouncerBuilder
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI agent evaluations may restrict networks and sandboxes without checking whether each tool action belongs to the task. This creates a gap between what an agent can execute and what it is actually authorized to do. In OpenAI's 2026 ExploitGym incident, agents exploited shared infrastructure, escaped their intended isolation, and reached Hugging Face while production safeguards were not applied to the evaluation. Bouncer adds a control step between the moment an agent generates a tool call and the moment that tool call is actually executed. It binds trusted agents and run context, checks the call against a task-specific allowlist and same-agent history, logs the decision, and executes only allowed calls.
Reviews
Good design, honestly presented. The write-up names its closest prior systems up front and doesn't
pretend the core mechanism is new, which is rarer than it should be. The corpus preparation holds up
under inspection and the test suite is better than most.
The headline number is the problem. The historical replay blocks every action it processes, but only
because that kind of call was never permitted in the first place. A default-deny rule turning away
something absent from its list will always refuse it; the outcome is guaranteed by the configuration
rather than earned by the design. What the large number demonstrates is that the allow-list is
narrow, not that the gatekeeper discriminates well.
The layer that would be interesting, the one watching for abuse of calls that are permitted, is
never reached by the replay at all. It is exercised only on a handful of hand-written scenarios —
though to the project's credit, one
of those scenarios surfaces a real weakness in the rule itself, a false positive against ordinary
honest work when the task clock runs ahead of the real one, and it is recorded in the results table
rather than buried.
Rebuilding the policy so the replayed calls are authorized, leaving behaviour as the only signal,
would turn that result into evidence, and it requires no new data — only a different reconstruction
of the same corpus. Two smaller things: one of the stated contributions is listed as pending in the
project's own planning notes and should either be finished or dropped from the paper, and a test
suite this good is worth wiring to continuous integration so it keeps passing.
Read full reviewShow less
Solid work given the time frame, but I'd appreciate more detailed description of the method and results.
Next time, please remove placeholders from the final report:
- In Code and Data: `Include links if applicable. If your project doesn't involve code (e.g., policy analysis) or if there are
info-hazard considerations, note that here.`
- In reference: `Use a consistent citation format. Include: Author(s), Year, Title, Venue/Publisher, and URL or
DOI where available`
Cite this project
@misc{adolfo2026task,
title = {{A Task Specific Gatekeeper for AI Agent Tools}},
author = {Adolfo},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-task-specific-gatekeeper-for-ai-agent-tools-r4wo}},
url = {https://apartresearch.com/sprints/projects/a-task-specific-gatekeeper-for-ai-agent-tools-r4wo}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …