When the Evaluator Becomes the Attack Surface: Security Risks in Agentic AI Evaluation
Saleha Muzammil · Team Catwoman
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project audits the published deployment configurations of 25 agentic AI benchmarks to demonstrate that current security practices overlook the scoring path, an inbound channel where evaluated-agent outputs can reach privileged evaluation infrastructure, grading components, and live operator credentials. By analyzing 4,378 files across seven security variables, the research uncovers recurring exposure, including a traced path from agent output to a tool-using judge with unrestricted network egress, which existing automated security scanners fail to detect due to the relational nature of the vulnerabilities. To address this, the project proposes the Evaluation Deployment Contract, a machine-readable declaration and linter designed to enforce and verify safety assumptions between AI agents and their evaluation environments.
Reviews
This project audits 25 agent benchmarks to identify cases where the scoring infrastructure may expose credentials, answer keys, or other resources to evaluated agents, potentially allowing them to cheat.
However, the novelty is limited, as similar issues have already been studied in works such as BenchJack and HackDetect.
The proposed Evaluation Deployment Contract is the most distinctive contribution, but it is essentially a structured declaration with a basic checker that validates fields and whether cited code locations exist, without verifying that the evidence actually supports the security claims.
Therefore, its practical value and methodology remain unclear.
This is a strong security contribution. Treating the grader and scoring path as its own privileged attack surface. The audit is also much broader than a single anecdote, and the scanner comparison makes the point clear Many of these problems depend on relationships between the agent, grader, credentials and answer key, so ordinary configuration scanners miss them.
I wld like to see a broader or randomly sampled corpus and safe end-to-end validation of some of the traced paths, while keeping the current distinction between exposure and demonstrated compromise.
The audit identifies a consequential security boundary between agent-controlled outputs and privileged scoring infrastructure, supported by a substantial cross-project configuration review. Its treatment of unknown configurations and proposed deployment contracts provide useful tools for independent evaluators.
Cite this project
@misc{muzammil2026evaluator,
title = {{When the Evaluator Becomes the Attack Surface: Security Risks in Agentic AI Evaluation}},
author = {Saleha Muzammil},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-the-evaluator-becomes-the-attack-surface-security-risks-in-agentic-ai-evaluation-9qk0}},
url = {https://apartresearch.com/sprints/projects/when-the-evaluator-becomes-the-attack-surface-security-risks-in-agentic-ai-evaluation-9qk0}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …