Limitations in existing regulatory mechanisms when applied to attacks by AI agents
Juan Jeronimo Manriquez, Tomás Cifuentes Clavijo · Team AGWatch
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Current reporting mechanisms with investigative or regulatory capabilities have deficiencies that make reporting any autonomous AI agent related incidents to date —including the Hugging Face and DSEwiki incidents— infeasible for third parties. This conclusion was reached after conducting an investigation into the two main western AI related legal regimes, with California's SB 53 and the EU’s AI Act being examined, locating three main fault points in these systems. These systems are confidential or do not publish a public record; and fail either by not accepting reports from non-providers, or by having a threshold for what qualifies as an incident that disqualifies the incidents that have already happened. An exploration of the existing literature related to these incidents corroborated that these holes are real and significant. Two main artifacts were created. A complementary website that helps visualize where these systems fail, and a list of proposals that would amend the systems and strengthen them against future attacks carried out by AI agents.
Reviews
This project examines the recent incidents against the incident reporting criteria for the EU AI Office and California's SB53 to assess whether they fall within scope of reportable incidents and propose modifications to these criteria. This is a useful activity which could be of value to these regulators and other policy makers considering putting in place similar reporting requirements.
I found the paper clear, read naturally and had a good level of detail. There were some decisions made in the paper around which dimensions were important which would have benefited from being better defended:
Why do third parties need to be able to report incidents, given the legal obligations fall on the AI developers and they would have the best access to the details surrounding the incident? The examples you give of the wiki’s administrator and a third party with access gave enough to imply reasoning here, but this could have been made more explicit and backed up with data.
e.g. Evidence points to both incidents being first identified by the compromised organisations (Huggingface and DSEwiki) - if they suspected an AI related attack but could not identify the source or did not get a response back from the developer, this supports the argument for these being able to report.
Also, if a third party such as METR observe a developer withholding information, can they report? Perhaps under whistleblowing provisions.
Given only the DSEwiki was reported to the EU AI Office, the jurisdictional scope of the incident was likely a scope determinant. Do you see this as a problem? That the regulator would then have to rely on public reporting sources. Even if not impacting directly the relevant jurisdiction, should incidents which inform new dangerous capabilities be reportable? This could point to a clearer gap what is enforceable with patchwork regulations too.
Challenges could discuss false reports and verification.
Read full reviewShow less
This submission has set out to build a tool to help victims of agentic attacks file reports. It found there was nowhere to file. The authors observe that SB 53 exempts deceptive model behaviour occurring during an evaluation designed to elicit it, which is the setting the Hugging Face incident arose in. Additionally, the authors did not stop at their own reading of the law but checked it against the California agency's own position and against the Commission's refusal to say which mechanism OpenAI filed under.
However, the claim is broader than the evidence. Only the AI Act and SB 53 were examined, so what the paper can support is that no mapped AI-specific channel takes third-party reports. A compromised wiki administrator (a role the paper uses) still has breach notification, national incident channels and ordinary computer-misuse law, and the paper needs to explicitly say why those do not fill the gap. Article 73 of the EU AIA goes undiscussed. Second, the three failure points (threshold, who may file, no public record) are used as a framework but these are not justified.
Read full reviewShow less
Cite this project
@misc{manriquez2026limitations,
title = {{Limitations in existing regulatory mechanisms when applied to attacks by AI agents}},
author = {Juan Jeronimo Manriquez and Tomás Cifuentes Clavijo},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/limitations-in-existing-regulatory-mechanisms-when-applied-to-attacks-by-ai-agents-5olx}},
url = {https://apartresearch.com/sprints/projects/limitations-in-existing-regulatory-mechanisms-when-applied-to-attacks-by-ai-agents-5olx}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …