Skip to content
Sprint projectSep 13, 2026Charlottesville

When the Evaluator Becomes the Attack Surface: Security Risks in Agentic AI Evaluation

Saleha Muzammil · Team Catwoman

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: When the Evaluator Becomes the Attack Surface: Security Risks in Agentic AI Evaluation

Code (opens in new tab)
Share

This project audits the published deployment configurations of 25 agentic AI benchmarks to demonstrate that current security practices overlook the scoring path, an inbound channel where evaluated-agent outputs can reach privileged evaluation infrastructure, grading components, and live operator credentials. By analyzing 4,378 files across seven security variables, the research uncovers recurring exposure, including a traced path from agent output to a tool-using judge with unrestricted network egress, which existing automated security scanners fail to detect due to the relational nature of the vulnerabilities. To address this, the project proposes the Evaluation Deployment Contract, a machine-readable declaration and linter designed to enforce and verify safety assumptions between AI agents and their evaluation environments.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project audits 25 agent benchmarks to identify cases where the scoring infrastructure may expose credentials, answer keys, or other resources to evaluated agents, potentially allowing them to cheat.

    However, the novelty is limited, as similar issues have already been studied in works such as BenchJack and HackDetect.

    The proposed Evaluation Deployment Contract is the most distinctive contribution, but it is essentially a structured declaration with a basic checker that validates fields and whether cited code locations exist, without verifying that the evidence actually supports the security claims.

    Therefore, its practical value and methodology remain unclear.

  2. This is a strong security contribution. Treating the grader and scoring path as its own privileged attack surface. The audit is also much broader than a single anecdote, and the scanner comparison makes the point clear Many of these problems depend on relationships between the agent, grader, credentials and answer key, so ordinary configuration scanners miss them.

    I wld like to see a broader or randomly sampled corpus and safe end-to-end validation of some of the traced paths, while keeping the current distinction between exposure and demonstrated compromise.

  3. The audit identifies a consequential security boundary between agent-controlled outputs and privileged scoring infrastructure, supported by a substantial cross-project configuration review. Its treatment of unknown configurations and proposed deployment contracts provide useful tools for independent evaluators.

Cite this project

@misc{muzammil2026evaluator,
  title = {{When the Evaluator Becomes the Attack Surface: Security Risks in Agentic AI Evaluation}},
  author = {Saleha Muzammil},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/when-the-evaluator-becomes-the-attack-surface-security-risks-in-agentic-ai-evaluation-9qk0}},
  url = {https://apartresearch.com/sprints/projects/when-the-evaluator-becomes-the-attack-surface-security-risks-in-agentic-ai-evaluation-9qk0}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026