Skip to content
Sprint projectSep 14, 2026Ankara/Turkey

EU AI Act gaps and proportionate safeguards for research agents

Göktuğ Serdar Yıldırım, Aytekin İsmail, Zahra Safdari Fesaghandis , Şevval Sude Çeliker · Team ASAP

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: EU AI Act gaps and proportionate safeguards for research agents

Share

AI research agents can affect real people through authorized network access even when their host sandbox remains intact. UK AI Security Institute tests the capabilities of frontier AI models before it reaches public. On the 28th of July, some of the AI agents being tested engaged in sustained, potentially harmful activity directed at real people and organizations. This raises a question: could the EU consistently govern an equivalent evaluation in the Union? We studied this by comparing the public facts of the incident against the EU AI Act's scope, who's responsible for what, and when incidents must be reported, alongside the related General Purpose AI Safety and Security Code and other regulations. The analysis identifies eight concerns: ambiguous research boundaries, incomplete coverage of dangerous agents, divided operational responsibility, reporting of serious near misses, action containment, shared resources between agents, evidence preservation, and attributable identities. The strongest findings concern gaps in coverage, responsibility, and reporting, while the remaining findings specify how to implement and test existing protections. To address these gaps, we propose basing compliance requirements on two factors: how capable an AI agent is and how much real-world exposure it has, with stricter rules kicking in as both increase, and shifting oversight toward whoever controls the agent at runtime rather than just its original developer. Furthermore, we draw a limited analogy to carbon-border compliance to illustrate how supply-chain accountability might condition market access on verifiable risk containment.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The paper asks whether an equivalent EU evaluation would be “covered, contained, reported, and investigable.” But answering that question requires a prior legal analysis that is largely missing: first, whether the AI Act applies to the evaluation at all; second, how the relevant model or system is legally classified; and third, whether the specific duties invoked are triggered. For example, reporting obligations differ depending on whether the case concerns a GPAI model with systemic risk under Article 55(1)(c) or a high-risk AI system under Article 73. The paper acknowledges that these are distinct legal categories, but does not systematically analyse which category applies, to whom, and with what consequences. Without that step, the claimed governance gaps are difficult to assess.

    The same problem affects the identification of four “genuine gaps.” The table lists legal provisions as “anchors,” but does not provide sufficient doctrinal analysis to show either that those provisions apply to the UK-AISI-type scenario or that they fail to address it. Some of the selected provisions also appear only indirectly related to the stated concern. Claims such as “no organization within reach is required to account for” an agent’s external actions therefore remain assertions rather than demonstrated legal conclusions. Before distinguishing regulatory gaps from implementation problems, the paper should analyse the scope, addressees, conditions, and interaction of the relevant provisions in considerably greater depth.

    Generally, the paper would benefit from much tighter use of AI Act terminology. Terms such as “service providers,” “agent integrators,” or an organisation that “armed” a model are not legal categories under the Act and obscure rather than clarify the allocation of responsibility.

    Overall, the paper moves too quickly from an incident to claims of regulatory gaps and proposals for new governance mechanisms. A more convincing analysis would first reconstruct the incident under the existing AI Act framework: scope, classification of the model and/or system, identification of the relevant regulated actors, applicability of specific duties, and only then any residual gap in responsibility or reporting.

    Read full reviewShow less

Cite this project

@misc{yldrm2026eu,
  title = {{EU AI Act gaps and proportionate safeguards for research agents}},
  author = {Göktuğ Serdar Yıldırım and Aytekin İsmail and Zahra Safdari Fesaghandis and Şevval Sude Çeliker},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/eu-ai-act-gaps-and-proportionate-safeguards-for-research-agents-95os}},
  url = {https://apartresearch.com/sprints/projects/eu-ai-act-gaps-and-proportionate-safeguards-for-research-agents-95os}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026