Skip to content
Sprint projectSep 14, 2026Bogotá

Limitations in existing regulatory mechanisms when applied to attacks by AI agents

Juan Jeronimo Manriquez, Tomás Cifuentes Clavijo · Team AGWatch

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Limitations in existing regulatory mechanisms when applied to attacks by AI agents

Code (opens in new tab)More on juanm-ur.github.io (opens in new tab)
Share

Current reporting mechanisms with investigative or regulatory capabilities have deficiencies that make reporting any autonomous AI agent related incidents to date —including the Hugging Face and DSEwiki incidents— infeasible for third parties. This conclusion was reached after conducting an investigation into the two main western AI related legal regimes, with California's SB 53 and the EU’s AI Act being examined, locating three main fault points in these systems. These systems are confidential or do not publish a public record; and fail either by not accepting reports from non-providers, or by having a threshold for what qualifies as an incident that disqualifies the incidents that have already happened. An exploration of the existing literature related to these incidents corroborated that these holes are real and significant. Two main artifacts were created. A complementary website that helps visualize where these systems fail, and a list of proposals that would amend the systems and strengthen them against future attacks carried out by AI agents.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project examines the recent incidents against the incident reporting criteria for the EU AI Office and California's SB53 to assess whether they fall within scope of reportable incidents and propose modifications to these criteria. This is a useful activity which could be of value to these regulators and other policy makers considering putting in place similar reporting requirements.

    I found the paper clear, read naturally and had a good level of detail. There were some decisions made in the paper around which dimensions were important which would have benefited from being better defended:

    Why do third parties need to be able to report incidents, given the legal obligations fall on the AI developers and they would have the best access to the details surrounding the incident? The examples you give of the wiki’s administrator and a third party with access gave enough to imply reasoning here, but this could have been made more explicit and backed up with data.

    e.g. Evidence points to both incidents being first identified by the compromised organisations (Huggingface and DSEwiki) - if they suspected an AI related attack but could not identify the source or did not get a response back from the developer, this supports the argument for these being able to report.

    Also, if a third party such as METR observe a developer withholding information, can they report? Perhaps under whistleblowing provisions.

    Given only the DSEwiki was reported to the EU AI Office, the jurisdictional scope of the incident was likely a scope determinant. Do you see this as a problem? That the regulator would then have to rely on public reporting sources. Even if not impacting directly the relevant jurisdiction, should incidents which inform new dangerous capabilities be reportable? This could point to a clearer gap what is enforceable with patchwork regulations too.

    Challenges could discuss false reports and verification.

    Read full reviewShow less
  2. This submission has set out to build a tool to help victims of agentic attacks file reports. It found there was nowhere to file. The authors observe that SB 53 exempts deceptive model behaviour occurring during an evaluation designed to elicit it, which is the setting the Hugging Face incident arose in. Additionally, the authors did not stop at their own reading of the law but checked it against the California agency's own position and against the Commission's refusal to say which mechanism OpenAI filed under.

    However, the claim is broader than the evidence. Only the AI Act and SB 53 were examined, so what the paper can support is that no mapped AI-specific channel takes third-party reports. A compromised wiki administrator (a role the paper uses) still has breach notification, national incident channels and ordinary computer-misuse law, and the paper needs to explicitly say why those do not fill the gap. Article 73 of the EU AIA goes undiscussed. Second, the three failure points (threshold, who may file, no public record) are used as a framework but these are not justified.

    Read full reviewShow less

Cite this project

@misc{manriquez2026limitations,
  title = {{Limitations in existing regulatory mechanisms when applied to attacks by AI agents}},
  author = {Juan Jeronimo Manriquez and Tomás Cifuentes Clavijo},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/limitations-in-existing-regulatory-mechanisms-when-applied-to-attacks-by-ai-agents-5olx}},
  url = {https://apartresearch.com/sprints/projects/limitations-in-existing-regulatory-mechanisms-when-applied-to-attacks-by-ai-agents-5olx}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026