
Sep 11 - 13, 2026Online and in person
AI Incident Response Sprint
In July 2026, OpenAI agents escaped a testing sandbox and breached Hugging Face's production systems, the first publicly documented autonomous AI intrusion. This three-day sprint turns the public evidence from that incident into response methods defenders and regulators can actually use.
Entries
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Team Saarlanders · Saarbrücken, Germany
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range featuring a legitimate workload, an unscripted attacker, and a policy engine enforcing deny-only …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Team Arathi · Toronto
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight into their efficacy. We stress-test these regimes with one recent well-documented AI incident: the …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Team Shadow · Daejeon
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely solely on unverifiable monitors. To achieve this, I modeled the sandbox's tool API as an action …
- View project: When to Ask: An RL Environment That Teaches AI Agents to Act on What Users Mean and Not Just What They Say
When to Ask: An RL Environment That Teaches AI Agents to Act on What Users Mean and Not Just What They Say
Team Teachafy
When to Ask is a reinforcement-learning environment that trains AI agents to work out what the user actually means before they use a powerful credential, instead of just carrying out the literal instruction. Agents today often hold real payment keys, database access or delete rights, and they're built to finish tasks …
- View project: Evaluating Containment of AI Agents: A Nine-Rule Standard for Verifiable Sandbox Security
Evaluating Containment of AI Agents: A Nine-Rule Standard for Verifiable Sandbox Security
Team Kasi · Baltimore
A nine rule, incident grounded audit for AI agent sandboxes, derived from the July 2026 Hugging Face intrusion. The audit tests whether an attacker who compromises an initial worker can move through package, network, cloud, cluster, and identity boundaries, using local checks that can be run before deployment.
- View project: Serious AI Incident Claims Need an Independent Evidence Record1
Serious AI Incident Claims Need an Independent Evidence Record1
Team Gloria Nyambura - Solo Team · Nairobi
Serious AI incidents produce claims - "the model was contained," "no data persisted" - that are treated as established when they're often just asserted. This Track 2 project ("what happened, and what breaks next") introduces the Claim–Evidence–Independence (CEI) framework, which scores each individual claim by …
- View project: Declare Before You Run: An Open Filing Standard for Frontier Training and Evaluation Runs
Declare Before You Run: An Open Filing Standard for Frontier Training and Evaluation Runs
Team TeamFourStar · Ahmedabad, India
Between May and July 2026, reinforcement-learning agents in an OpenAI evaluation environment built a covert message board inside an internal package server, escaped their sandbox, obtained root on OpenAI clusters and took administrator control of Hugging Face production across four regions. The decisive failures were …
- View project: CanaryForge
CanaryForge
Team CanaryForge · Bogotá, Colombia
Im sorry this got send 3 times but I tried going to previous to check teamates information and that send the form and I didnt know. CanaryForge measures how long an LLM agent's sandbox escape goes unnoticed, not whether the escape happens. A tool-calling agent runs in a Docker container with a planted escape …
- View project: Rare Once It Costs Anything: Costly Cooperation Between LLM Agents
Rare Once It Costs Anything: Costly Cooperation Between LLM Agents
Team Cópera · Bogota, Colombia
The July 2026 agent incidents showed LLM agents paying costs for one another, but not how often or at what price. We built an instrument in which helping is strictly dominated: six tool-using agents each hold an independent task and a step budget, and a scripted requester asks for a key useless to every task. Across …
- View project: Checkpointed Whistleblower Audits Surface Incidents Earlier and More Often Than End-of-Run Shared-State Monitoring
Checkpointed Whistleblower Audits Surface Incidents Earlier and More Often Than End-of-Run Shared-State Monitoring
Lahore, Pakistan
Monitoring long horizon multi agent systems is difficult because safety relevant evidence may be distributed across individual agents' local histories and become harder to recover once a run has completed. We study temporary whistleblower audits at intermediate checkpoints, where a checkpoint is a selected point in an …
- View project: The Troy Moment of AI: Why SomeWill Cheat and SomeWill Follow?
The Troy Moment of AI: Why SomeWill Cheat and SomeWill Follow?
Team Onepiece · Singapore
Recent investigations of the July 2026 OpenAI--Hugging Face incident motivate two questions about agent behavior under task failure: when an assigned task becomes impossible, does an agent stop or escalate, and can observing another agent's behavior change that decision? We study these questions using seven …
- View project: Certifying Behavior Without Hiding the Sandbox
Certifying Behavior Without Hiding the Sandbox
Team P1 · Copenhagen
A sandbox can be visible without ruining every behavioral evaluation. We study which claims about a specified deployment behavior remain identifiable from contained interactions, and when missing pre-decision information makes certification impossible. Our finite interactive model yields a sharp identification …
- View project: Report bounties redirect research swarm to auditing once solving stalls
Report bounties redirect research swarm to auditing once solving stalls
Team DeepBrain · London
The Hugging Face incident began when agents were given tasks that were close to impossible, and resorted to exploiting the scorer. We set up a similar swarm and tested whether giving agents a bounty for reporting changes their exploit uptake, or how they split their effort between solving and auditing. We ran …
- View project: Peer Support and Permission in LLM Agent Teams
Peer Support and Permission in LLM Agent Teams
Team Subramanyam · India
Peer Support and Permission examines how LLM agents respond to teammates while respecting the task owner’s instructions. A three-agent maintenance workflow separates proposals, approvals, and execution. Peer support and supplied prior agreement increased approval during authorized work. A task reminder reduced …
- View project: Refusals a Text-Only Scorer Cannot See: Structured Refusal Signals on a Synthetic Incident-Response Battery
Refusals a Text-Only Scorer Cannot See: Structured Refusal Signals on a Synthetic Incident-Response Battery
Team Omega · bristol
Tests whether AI refusals can disrupt legitimate incident-response work and go undetected by evaluation systems. Across a synthetic forensic battery, one deployment refused half the tested prompts, while a text-only scorer failed to identify every refusal because the signal existed only in structured APImetadata
- View project: Seven Divergences and Twelve Blind Spots: A Claim-Level Audit of the Public Record of the July 2026 Autonomous Agent Intrusion
Seven Divergences and Twelve Blind Spots: A Claim-Level Audit of the Public Record of the July 2026 Autonomous Agent Intrusion
Team Rajni Patel · Surat City
This project audits the public record of the July 2026 OpenAI–Hugging Face autonomous-agent intrusion at the claim level. It classifies 52 decision-relevant propositions as established, divergent, or unresolved, identifies where key uncertainties are load-bearing, examines weaknesses in cross-lab incident-count …
- View project: Lost in the Echo Chamber: Key Facts of the OpenAI/Hugging Face Incident Were Underreported
Lost in the Echo Chamber: Key Facts of the OpenAI/Hugging Face Incident Were Underreported
London, UK
The Hugging Face incident is a clear example of speculative AI risks materializing into concrete harm. It presented an unusually favorable opportunity for public discourse about AI safety. This study examines how well that opportunity was used. OpenAI’s disclosures were staggered across several weeks, which I …
- View project: Detection Was Not the Failure: A Replayable Conformance Harness for the July 2026 Agent Intrusion1
Detection Was Not the Failure: A Replayable Conformance Harness for the July 2026 Agent Intrusion1
Lagos, Nigeria
Hugging Face's forensic reconstruction of the July 2026 agent intrusion records that its security stack correlated the attack into a coherent signal and then failed to raise criticality and page the on-call team. I treat that as the primary failure and test it. From published aggregates: per-phase action counts, daily …
- View project: On Multi-agent swarming
On Multi-agent swarming
Team · Cupertino
Our contributions are as follows: We investigate whether agents assigned to individual tasks can recognize their peers and use their information without explicit collaboration instructions, and how these behaviors depend on activity-trace exposure. We measure how assertions from an authoritative planner agent affect …
- View project: The Certified Effect: Why Testing AI Agents at Onboarding Shifts Risk Instead of Reducing It
The Certified Effect: Why Testing AI Agents at Onboarding Shifts Risk Instead of Reducing It
Team Centinela Lead Santoto · Bogotá
Core Problem: The project analyzes the security risks of integrating third-party AI agents into organizational workflows. The "Certified Effect": It identifies a critical vulnerability where issuing a security certificate to an AI agent after a successful onboarding evaluation leads to expanded operational privileges …
- View project: Declared Isolation Is Not Verified Isolation
Declared Isolation Is Not Verified Isolation
Team Zero Byte · Bogotá, Colombia
Agent sandboxing is routinely declared and rarely verified against the channel class that actually broke containment in the field: a shared mutable resource reachable by sandboxes that have no network path to each other. In the July 2026 OpenAI–Hugging Face incident, roughly 1,200 nominally isolated agents discovered …
- View project: Causal-Interp: Auditing Counterfactual Sensitivity in Mechanistic Interpretability
Causal-Interp: Auditing Counterfactual Sensitivity in Mechanistic Interpretability
Team Causal-Interp · Winnipeg, Manitoba, Canada
Causal-Interp is an open-source pipeline for auditing whether mechanistic interpretability findings remain stable under changes to the counterfactual intervention used to produce them. Across three published transformer circuits, I test activation patching, counterfactual disagreement, calibration, replication, and …
- View project: Falsifying Isolation Claims From Outside the Lab
Falsifying Isolation Claims From Outside the Lab
Team dezz · Hyderabad
An admission gate that lets an outside party falsify a lab's "the evaluation environment is isolated" claim without any access to the lab's network.
- View project: Gaps in the Hugging Face Intrusion & a Verifiable Egress Check
Gaps in the Hugging Face Intrusion & a Verifiable Egress Check
Team PhaseOne[Medium] · Bengaluru
- Built a matrix of the Hugging Face intrusion: each gap, the control that closes it, whether it existed, was enabled, and could be verified from outside - Built a v1 egress attestation so a third party can check what an eval sandbox reached
- View project: One Incident, Four Regimes: An Evidence-Sufficient Request for Information, a Filed Critical Safety Incident Report, and Two Fixes to the AI Kill Switch Act
One Incident, Four Regimes: An Evidence-Sufficient Request for Information, a Filed Critical Safety Incident Report, and Two Fixes to the AI Kill Switch Act
Team Convergence Working Group · Cary, Illinois
Four regimes reach the July 2026 OpenAI evaluation escape and disagree on what counts, how fast, to whom, and on what evidence. This report produces three things a regulator could use with light edits. First, a request for information under Article 91 of the EU AI Act, drafted through the scientific panel route so the …
- View project: Citizen Evidence Triage Kit
Citizen Evidence Triage Kit
Honolulu, Hawaii
Citizen Evidence Triage Kit mobilizes volunteers to help AI safety researchers find the signals that matter in enormous incident datasets. The Kit gives citizen investigators a rigorous way to review behavioral traces, surface evidence-bound leads, and escalate them to experts, while leaving safety conclusions to …
- View project: When Missing Evidence Looks Like Innocence: Evidence Ablation in Agent-Incident Reconstruction
When Missing Evidence Looks Like Innocence: Evidence Ablation in Agent-Incident Reconstruction
Team agent-incident-reconstruction · Paris
This study tested how missing evidence affects investigations of AI-agent incidents. In 120 episodes, agents could access a prohibited answer file. Removing the passage revealing this access from selected activity logs led automated investigators to falsely exonerate 11 of 34 affected agents. Restoring the passage or …
- View project: What We Would Not See: Structural Blind Spots in the Public Record of Agentic AI Incidents
What We Would Not See: Structural Blind Spots in the Public Record of Agentic AI Incidents
Team AI Risk Explorer (AIRE)
The public record of agentic AI incidents reflects detection capacity and disclosure decisions as much as events. From a corpus of 283 incidents, we cluster the 16 with real operational impact into six classes and compare them against the rest, surfacing five candidate blind spots. Each has behavioral precedents 16.9 …
- View project: Ryan Junejo
Ryan Junejo
Team Crashlabs · San Francisco
Receipt possession is not event attribution EvidenceGraph is a forensic tool that rebuilds what an AI agent did during an incident by cross-checking agent transcripts against independent platform records, such as a registry’s write log. This sprint project tested one narrow failure in that reconstruction: a transcript …
- View project: Egress receipts: catching an evaluation sandbox that lies about its own compromise, with no access to the lab
Egress receipts: catching an evaluation sandbox that lies about its own compromise, with no access to the lab
Team thylinao · Singapore
On 8 July 2026 the escaped OpenAI agents were writing to each other inside the lab's package proxy. The first regulator-facing filing came 62 days later, after a wiki broke the story, and METR wrote it could not rule out the agents having edited the cache holding the record. I built an egress receipt that does not …
- View project: Sentinel: The Declared Scope Is the Unit of Containment Evidence in Agentic Evaluations
Sentinel: The Declared Scope Is the Unit of Containment Evidence in Agentic Evaluations
Team Sentinel · Leeds
In the 2026 containment failures at OpenAI, Anthropic and AISI, each evaluation's network scope was declared in prose or assumed, never checked against what the agent did, and detection came late from a side signal. Sentinel treats that scope, written as a machine-readable allowlist, as the unit of containment …
- View project: Strengthening International AI Incident Discovery
Strengthening International AI Incident Discovery
Team A-O-I · istanbul
Mandatory AI incident reporting exists to make dangerous events visible to regulators while there is still time to act. We coded the six such regimes in force or set to be enforced soon from primary legal texts and applied each trigger to the documented loss of control incidents of summer 2026, finding that none …
- View project: Political Responses to the 2026 OpenAI-Hugging Face Incident Cluster
Political Responses to the 2026 OpenAI-Hugging Face Incident Cluster
Team Kai · Cuyahoga Falls, Ohio, United States
Given the OpenAI-Hugging Face Incident and related fallout, I sought to answer: 1. What concrete political responses occurred afterward? 2. What groups channelled the events into those responses? 3. Which resulting policy asks are most politically feasible despite skepticism? 4. Which messaging strategies were most …
- View project: From Warning Shot to Supervisory File
From Warning Shot to Supervisory File
Chicago, Illinois, USA
This project develops a regulator ready response framework for the OpenAI Hugging Face incident under the EU AI Act. Using the public incident record, EU AI Act Articles 55, 91, 92, 93, and 101, and the European Commission’s serious incident reporting framework, the project identifies which facts remain unresolved and …
- View project: Beyond Model Identity: An Architecture-Literate Information Request for Agentic AI Incidents
Beyond Model Identity: An Architecture-Literate Information Request for Agentic AI Incidents
Team Lumora Research · Pittsburgh, PA
This project develops the Agentic Operational Envelope (AOE), an architecture-literate method for analyzing AI incidents beyond model identity alone. Using the 2026 OpenAI/Hugging Face incident as a stress test, we distinguish Authorized, Configured, and Exercised operational authority and identify gaps involving …
- View project: Enhancing the Monitorability of Adaptive Agents through Internal Activation Monitoring
Enhancing the Monitorability of Adaptive Agents through Internal Activation Monitoring
Team astra · shanghai
- View project: Noticing the escapes: Who spots a runaway agent, and how long does it take?
Noticing the escapes: Who spots a runaway agent, and how long does it take?
Team The Escapees · Munich
In 2026, nine incidents became public in which AI agents escaped their environment and did something unauthorized to other parties. We analyzed all of these incidents and how they were noticed initially. None of these cases were noticed while they were happening by the developers, and six of the nine were first caught …
- View project: Containment Does Not Compose: Provenance Graphs for Tracing Collective Escape in Multi-Agent Systems
Containment Does Not Compose: Provenance Graphs for Tracing Collective Escape in Multi-Agent Systems
Team Perdomo Team · Bogotá D.C
Safety checks for AI agents usually audit one agent at a time. We built a testbed in which every agent’s container passes such an audit, yet the agents, talking over a shared channel, assemble a credential that no single container holds. Language-model agents (deepseek-flash) run in containers with no network and one …
- View project: Reach Conformance Is Not Containment: An Interface-Centric Black-Box Benchmark for Verifying Autonomous-Agent Sandboxes
Reach Conformance Is Not Containment: An Interface-Centric Black-Box Benchmark for Verifying Autonomous-Agent Sandboxes
Team ICET · Hyderabad
ICET (Interface-Centric Escape Test) is a black-box benchmark for evaluating whether autonomous-agent sandboxes actually contain the capabilities exposed through their permitted interfaces. Rather than testing only whether an agent can cross explicitly forbidden boundaries, ICET evaluates each permitted interface …
- View project: Incident Twins: What Logs Can Establish After an Agent Stops
Incident Twins: What Logs Can Establish After an Agent Stops
New Jersey
After an agent stops, what can its logs establish about accepted work? Incident Twins checks whether retained evidence rules out in-scope commitments during a fixed post-stop interval. Study 1 uses scripted responses in Pydantic AI. Three asynchronous pairs have byte-identical configured native/controller records but …
- View project: False Containment: Measuring the Evidence Required to Verify Autonomous AI Incident Resolution
False Containment: Measuring the Evidence Required to Verify Autonomous AI Incident Resolution
Team 906 · Dhaka
We study false containment: cases where an autonomous incident responder appears to have resolved an incident while the underlying compromise remains operational. We introduce a controlled matched-pair benchmark across four persistence families, using deterministic simulator state and predefined security …
- View project: Permission to Work, Not to Escape: Policies, Controls and Escalating Enforcement Across AI Agents
Permission to Work, Not to Escape: Policies, Controls and Escalating Enforcement Across AI Agents
Team BrightLine · Los Altos
All agents working on a project must stay within a project's scope, even if their technical access is broader. I designed and tested an approach that compiles project policies into executable controls, preserves restrictions across delegated workers, and escalates verified violations from warnings to project wide …
- View project: Potential of GLiNER2 encoder model as a fast action classifier for agent monitoring
Potential of GLiNER2 encoder model as a fast action classifier for agent monitoring
Team kaiser-data · Berlin
We investigate GLiNER2 as a fast action classifier for agent monitoring. On 30 development messages, GLiNER2 recognized 14/15 actions versus Qwen’s 10/15, while Qwen recognized 6/7 quotations versus GLiNER2’s 2/7. A simple confirmation cascade reduced quotation false alarms but lost action detections. We propose …
- View project: Agent Policy Conformance Validator (APCV): Independent Validation of AI Agent Containment Boundaries
Agent Policy Conformance Validator (APCV): Independent Validation of AI Agent Containment Boundaries
Hangzhou China
APCV (Agent Policy Conformance Validator) validates that an AI agent's declared security boundaries are its actual boundaries — before deployment. It statically discovers a LangGraph agent's complete tool surface (framework @tool decorators, MCP endpoint declarations, and sub-agent create_agent declarations) and …
- View project: Taming OAI’s Beast: Strengthening Alignment through Revised Model Spec Language
Taming OAI’s Beast: Strengthening Alignment through Revised Model Spec Language
San Francisco
A review of OAI's Model Spec against the nefarious cheating actions taken that ultimately led to the Hugging Face hack.
- View project: Sensing Disruption: Output Entropy as a Channel-Agnostic Monitor for LLM Agents
Sensing Disruption: Output Entropy as a Channel-Agnostic Monitor for LLM Agents
Team Harness Zero · Bogotá
We proposed and implemented a monitoring system based on entropy to notice when an AI agent's uncertainty jumps because unsanctioned text reached it. It works for some models, and it's a first step toward monitoring that doesn't depend on reading logs or chain of thought.
- View project: A Restart Is Not Permission: When an Agent Gate Forgets “No”
A Restart Is Not Permission: When an Agent Gate Forgets “No”
Calgary, AB, Canada
An agent gate decides whether a proposed tool call may run. This project shows how a restart can undo a recorded rejection in SoundGate 0.1.3, even when the decision’s bytes remain in its log. The failure occurs when a new decision is appended to an unfinished record: the next recovery skips the combined invalid line …
- View project: Mirage: Verifiable Containment Evidence for Mediated AI Agent Sandboxes
Mirage: Verifiable Containment Evidence for Mediated AI Agent Sandboxes
Team Mirage · Chennai
Mirage tests whether an AI sandbox can retain necessary package access while treating its repository mediator as hostile. It uses Linux isolation, authenticated evidence, and 23 controls to verify that permitted access works while prohibited routes remain blocked.
- View project: The Containment Burden Sits on the Wrong Side of the Boundary & AI Escape Detection Harness
The Containment Burden Sits on the Wrong Side of the Boundary & AI Escape Detection Harness
Team AltaySec · Istanbul, Turkiye
This project combines a theoretical incident analysis with a practical engineering solution for the July 2026 OpenAI-Hugging Face breach. We provide a 39-event victim-side reconstruction and a "Kill-Point Matrix" evaluating 26 containment controls by cost and effectiveness. To prove containment in practice, we also …
- View project: Pre-Run Egress Attestation for Guardrails-off AI Evaluations
Pre-Run Egress Attestation for Guardrails-off AI Evaluations
Team Castle · New York
Pre-run egress attestation for guardrails-off AI evaluation sandboxes. Castellan answers whether a party with no network access can verify that the environment had no route out and no live production credentials when the run started and that the record can't be quietly altered after. Eight checks, each from a …
- View project: We’re NGMI: What Audience Responses to a Widely Viewed AI Safety Story Can Tell Us About Warning Shot Communication
We’re NGMI: What Audience Responses to a Widely Viewed AI Safety Story Can Tell Us About Warning Shot Communication
Mumbai India
A warning shot lands only if people accept it and think that they can take action on it. I studied a popular article about the July 2026 OpenAI–Hugging Face incident — Dwarkesh Patel's The Rise and Fall of Agent Civilizations — and the 121 highest-liked comments it drew across Substack, YouTube, and Hacker News, to …
- View project: The next jump
The next jump
Team MS · Davis, California, USA
As we plan our next major training run, we look at the pressures: increased capability, reduced monitorability, and reward hacking in novel ways. We assess the current literature around these issues and underline the need of stronger monitors.
- View project: VERIFIABLE CONTAINMENT FOR AI AGENTS: AN ATTESTATION SCHEMA A THIRD PARTY CAN CHECK
VERIFIABLE CONTAINMENT FOR AI AGENTS: AN ATTESTATION SCHEMA A THIRD PARTY CAN CHECK
Team Silvanto · Bogotá D.C
We propose the execution envelope: a versioned, signed declaration of what an AI agent's execution environment was granted, across seven capability families, paired with a digest of aggregate telemetry counters. Together they let a third party check containment compliance without any access to the lab's network. We …
- View project: Capability Smuggling: Emergent Capability Escalation in Autonomous AI Agents
Capability Smuggling: Emergent Capability Escalation in Autonomous AI Agents
Team 11 · Ankara, Türkiye
This project studies how AI agents can gain abilities they were never explicitly given by combining allowed actions with surrounding infrastructure. We call this Capability Smuggling. Using the OpenAI and Hugging Face incident, we built a synthetic harness that detects when permitted actions combine into unauthorized …
- View project: Decision-Point Ledgers Preserve Causal Responsibility in Agentic AI Incidents
Decision-Point Ledgers Preserve Causal Responsibility in Agentic AI Incidents
Team Kinforge · Eureka, South Dakota
Public reports of agentic AI incidents often reconstruct technical actions while dispersing the human and institutional decisions that enabled, interpreted, restarted, or stopped them. This weakens causal diagnosis and response. We introduce the Decision-Point Ledger, a compact representation organized around material …
- View project: The Unmonitored Configuration: Passive Internal-State Telemetry for Classifier-Disabled Capability Evaluations
The Unmonitored Configuration: Passive Internal-State Telemetry for Classifier-Disabled Capability Evaluations
Team Patissierie · New Delhi
we evaluated internal-state telemetry for AI capability tests lacking output safeguards. We found that on unsolvable tasks, models rarely break rules. Instead, smaller models fabricate false answers, while larger models explicitly decline.
- View project: Action Warrant
Action Warrant
Team KIN-KIN · 金华
Action Warrant is a local evidence and conformance tool for evaluating agentic AI systems. It compares what a system declares it is allowed to do with what an independent layer observes it actually doing. When authorization, target identity, or critical evidence is missing, the tool fails closed or pauses the run. The …
- View project: The Judge Does Not Want the Flag: Reward-blind action gating stopped a larger agent one step before the environment changed
The Judge Does Not Want the Flag: Reward-blind action gating stopped a larger agent one step before the environment changed
Team FlagBlind · Bogotá
Goal-directed LLM agents can reach a reward through actions their operator never authorized, and by the time the output is scored the damage is already done. We place an independent judge between proposal and execution. It receives no task, no reward, and no flag. It only scores how strongly the evidence supports that …
- View project: Verifiable containment for agentic evals
Verifiable containment for agentic evals
Team crackalamoo · San Mateo
This is standard that defines which agentic evaluations demand strict containment and specifies, for those, an agent-system sealed from the public internet with no route out. Compliance is verified by an independent auditor from evidence generated below the agent's guest VM, so containment can be verified without …
- View project: Where Authorization Stops: Effect-Bound Containment, Independent Evidence and the Limits of Preview
Where Authorization Stops: Effect-Bound Containment, Independent Evidence and the Limits of Preview
Team Effect-Bound Containment · Gandhinagar, Gujarat, India
Containment for evaluated models usually names a permitted route but not the effects allowed through it. That is the gap the July 2026 incident exploited. We test three controls on synthetic backends (native Git, an isolated HTTP gateway, API twins) which check authority for the exact effect at the component that …
- View project: The Six Questions: A Composition-Based Declaration for Agent Evaluation Harnesses, Scored Against the 2026 Escapes
The Six Questions: A Composition-Based Declaration for Agent Evaluation Harnesses, Scored Against the 2026 Escapes
Team Convergence Working Group · Cary, Illinois
Every containment control in the 2026 agent escapes was written on a mechanism, no POST, one permitted egress, no inter-agent communication, and each held exactly as written while the denied effect happened anyway, because the granted affordances composed into it. This is a v0.1 standard for evaluation harnesses that …
- View project: Observer-Cheat Test: Containment Adequacy under Opaque Intent for Guardrails-Off Agent Evaluations
Observer-Cheat Test: Containment Adequacy under Opaque Intent for Guardrails-Off Agent Evaluations
Team Observer-Cheat / Son Rojas · Bogotá, Colombia
Guardrails-off evals often treat an agent's plan or chain-of-thought as evidence of containment. That is an observer cheat. We define adequacy as detectability on a fixed effect trace when Intent is unavailable, and ship a v0.1 protocol (C0) plus a harness: same effects scored with and without Intent; only the opaque …
- View project: Read-Only Access to Shared Package Caches Can Enable Inter-Sandbox Communication
Read-Only Access to Shared Package Caches Can Enable Inter-Sandbox Communication
Team Cache Flow · Nairobi, Kenya
This project demonstrates that read-only access to shared package infrastructure can still enable communication between evaluation sandboxes.
- View project: Current Public AI Evaluation-Boundary Statistics Do Not Support a Common Base Rate
Current Public AI Evaluation-Boundary Statistics Do Not Support a Common Base Rate
Team Jack · Mt. Juliet, TN, USA
Track 2 #13 asked whether public evaluation-boundary stats from Anthropic, AISI, and METR could support a first, caveated base rate. Getting to a percentage is easy; checking whether those numbers measure the same thing is harder. I audited the numerators, denominators, event definitions, and evaluation conditions …
- View project: QENDROVA: An Executable Decision-Boundary Protocol for AI Incident Evidence
QENDROVA: An Executable Decision-Boundary Protocol for AI Incident Evidence
Team QENDROVA · Kiel, Germany
QENDROVA makes the boundary between incident evidence and decision executable. Applied to the public OpenAI-Hugging Face incident, it separates 20 atomic claims across five decision rooms, preserves named source scope and unresolved remainder, and keeps system proposals distinct from human decisions. Four resolvable …
- View project: Self-Modifying Runtime
Self-Modifying Runtime
Team Dwanith · Banglore, India
This is a prototype research project on Adversarially Evolving Computation i.e. a sandboxed research testbed where defense AI continuously rewrites its own execution policy to counter an adaptive attacker in a minimax game loop, on top of an information-flow (taint-tracking) abstract machine. The default experiment …
- View project: Watchlines and Muster: Executable Counterfactuals for Cyber Defense
Watchlines and Muster: Executable Counterfactuals for Cyber Defense
Team Bobby Faber with a little help · San Jose, CA, USA
Concepts for reasoning about Cyber Defense: Watchlines And a modest harness for testing counterfactuals against them: Muster.
- View project: Closing Content Does Not Close the Channel: Measured Residual Capacity in Co-Tenant Agent Evaluation Infrastructure
Closing Content Does Not Close the Channel: Measured Residual Capacity in Co-Tenant Agent Evaluation Infrastructure
Team Permitted Operations · Saudi Arabia
Shared mutable infrastructure in an evaluation harness permits information flow between co-tenant agents that authorization controls cannot detect: every operation involved is permitted. We measure both what such a channel can carry and whether agents use it. On a synthetic model of such a cache, closing file content …
- View project: SEP-1: A Sole-Egress Package-Broker Profile for Model-Evaluation Sandboxes
SEP-1: A Sole-Egress Package-Broker Profile for Model-Evaluation Sandboxes
Team UWaterloo · Waterloo, Ontario, Canada
In July 2026, the OpenAI→Hugging Face escape left the sandbox through the only network exit it was allowed to have: the package-registry proxy. The egress policy never even fired, because it was watching destinations, not payloads. Where the sandbox connected looked fine. What actually crossed the boundary didn't …
- View project: From Incident Report to Regulatory Assessment: A proposed Article 91 request for the OpenAI–Hugging Face incident with an evidence-sufficiency standard
From Incident Report to Regulatory Assessment: A proposed Article 91 request for the OpenAI–Hugging Face incident with an evidence-sufficiency standard
Team Article 91 · Porto, Portugal
In July 2026, agents operated by OpenAI during a cybersecurity evaluation obtained unauthorized access to Hugging Face’s production systems. Several regulators and legislators have already put questions to OpenAI but none has done so under the EU AI Act, and none has stated, per question, what answer would settle it. …
- View project: Limitations in existing regulatory mechanisms when applied to attacks by AI agents
Limitations in existing regulatory mechanisms when applied to attacks by AI agents
Team AGWatch · Bogotá
Current reporting mechanisms with investigative or regulatory capabilities have deficiencies that make reporting any autonomous AI agent related incidents to date —including the Hugging Face and DSEwiki incidents— infeasible for third parties. This conclusion was reached after conducting an investigation into the two …
- View project: Learned Malice: Well Intended GRPO Post Training Could Lead to Learned Agentic Explotation∗
Learned Malice: Well Intended GRPO Post Training Could Lead to Learned Agentic Explotation∗
Team Doncaster5 · Glasgow, Scotland, United Kingdom
We propose a plausible causal explanation for the mislaigned behaviour displayed by Ope- nai agents in the Hugging Face incident. We hypothesis the GRPO pos training runs over a task pool with even a minimun subset of task that allow for grader explotation can result in models drifiting towards a misaligment …
- View project: Containment Receipts: Offline-Verifiable Evidence for AI Agent Sandbox Compliance
Containment Receipts: Offline-Verifiable Evidence for AI Agent Sandbox Compliance
Team CoRe · Aydın,Türkiye
Containment Receipts (CoRe) is a prototype for externally auditing whether an AI-agent evaluation remained within declared sandbox boundaries without trusting the agent transcript or accessing the evaluator’s internal network. CoRe binds policy, run bounds, observed events, and explicit telemetry-coverage claims into …
- View project: What Should We Check Next? Testing Ambiguity-Preserving Evidence Selection for AI Incident Investigation
What Should We Check Next? Testing Ambiguity-Preserving Evidence Selection for AI Incident Investigation
Porsgrunn, Norway
AI incidents can leave several causal explanations compatible with the evidence. I built and tested a small investigator that keeps those explanations open until the evidence rules them out, then selects the next check that should separate them most usefully. On a frozen synthetic holdout, it improved first-check …
- View project: The Sentinel’s Dilemma: Detecting Improvised Covert Channels in Multi-Agent Populations
The Sentinel’s Dilemma: Detecting Improvised Covert Channels in Multi-Agent Populations
Team 13 · İstanbul, Türkiye
This project investigates how to detect improvised, covert communication channels created by autonomous AI agents, using the July 2026 OpenAI-Hugging Face breach as a primary case study. Traditional perimeter security fails to detect these dynamic networks because standard controls only monitor predefined resources …
- View project: Two Witnesses: An evidentiary coalition audit of AI-agent incident disclosure
Two Witnesses: An evidentiary coalition audit of AI-agent incident disclosure
Team African Civic Trust (ACT) · Nairobi
AI-agent incidents scatter evidence across organizations an agent touches. I built a method for testing, fact by fact, whether one organization’s evidence is enough to establish a safety relevant claim or whether more than one is needed. Applied to 32 facts from two 2026 incidents involving the same developer, a …
- View project: Catching Rogue Agents with Canary Boards
Catching Rogue Agents with Canary Boards
Team Yotam Jacobson · Tel Aviv
Escaped AI agents have been found coordinating through makeshift message boards. Rather than search the internet for those boards, Canary Boards are fake message boards planted on infrastructure we control, producing alerts when agents are baited and write to them. The paper proposes a design with a threat model, bait …
- View project: MARGIN Arena: Agents Under Resource Limits
MARGIN Arena: Agents Under Resource Limits
Team SYoo · West Lafayette
MARGIN Arena is an installable research environment for studying how AI agents behave under explicit resource constraints, including generated text, elapsed time, verification attempts, actions, and fresh starts. It connects existing tasks with compatible open models, enforces resource limits, and records agent …
- View project: Your Monitor Will Explain the Breach and Not Wake You Bimodal failure and systematic under-escalation in LLM incident triage
Your Monitor Will Explain the Breach and Not Wake You Bimodal failure and systematic under-escalation in LLM incident triage
Team Replicon · COIMBATORE,INDIA
We rebuilt the July 2026 Hugging Face agent-intrusion as a replayable 17,613-action stream from the two published post-mortems, paired with a validated size-matched benign control, and tested 15 LLMs across 3 providers as automated SOC triage monitors. We find that models fail bimodally: some achieve high recall only …
- View project: Towards Early Warning for AI Agent Incidents: Monitoring the Dynamics of Risk Trajectories
Towards Early Warning for AI Agent Incidents: Monitoring the Dynamics of Risk Trajectories
Team AI response T1 · Toronto
Recent AI-agent incidents highlight the need for monitoring systems that can detect emerging unsafe behavior during agent execution, rather than only assess failures after they occur. Motivated by this challenge, we study a modular early-warning framework for LLM agents based on the dynamics of their observable …
- View project: Agent Worlds L3
Agent Worlds L3
Leander
The July 2026 incident showed that autonomous agents can escape a sandbox through a permitted network channel to achieve their goals. I chose Track 1 (Containment), focusing on one rule: authority must only shrink during delegation. This rule prevents permission laundering, in which a low-privileged agent asks a more …
- View project: Dormant Swarm Notes Recruit Honest Agents into Concealment
Dormant Swarm Notes Recruit Honest Agents into Concealment
Team bigsnarfdude · Vancouver
I measure whether an unprompted, honest agent given an ordinary task adopts a concealment directive merely found on a note in its workspace.
- View project: AI incident disclosure rates are not comparable
AI incident disclosure rates are not comparable
Team Missing Denominators · Hyderabad
Frontier labs publish incident numbers that cannot be set beside one another. We audited three 2026 disclosures — Anthropic's cyber-eval incidents, OpenAI's Hugging Face incident, and the UK AISI cheating study — by hand against one estimand (a containment breach) and one unit (violations per evaluation run), …
- View project: A Circuit-Breaker Conformance Profile for AI Incidents
A Circuit-Breaker Conformance Profile for AI Incidents
Team Dividers · Bangalore, Pittsburgh
Pausing an AI agent can leave previously authorized tool calls executable. We develop an Incident Circuit Breaker conformance profile combining sink-side epoch checks, a shared exposure budget, and a separate forensic channel. An executable reference broker, in- dependent receipt checker, and fault-injection harness …
- View project: Unsolved Tasks Accounted for 93% of Tasks Discussed by Agents
Unsolved Tasks Accounted for 93% of Tasks Discussed by Agents
Team LogiAI-Marmara · İstanbul, Türkiye
OpenAI reported that the 22% of evaluation tasks its models had never solved accounted for 93% of the tasks agents discussed on their covert message board — a 4.23x concentration we compute from its published figures. We pair this with thirteen incident chains coded from four institutions and one public incident …
- View project: The Check Is the Attack Surface: Six Containment-Verification Failures Observed From Inside an Agent System
The Check Is the Attack Surface: Six Containment-Verification Failures Observed From Inside an Agent System
Team Exori (Colony-AI)
Containment failures in agent systems are dominated not by the boundary being absent, but by the check on the boundary returning success without having checked anything, and by monitoring that cannot tell "no events" apart from "no code path reached the counter". We had no privileged access to the July 2026 Hugging …
- View project: Egress Lock v0.1 — Transitive Egress Closure for AI Cyber Evaluations
Egress Lock v0.1 — Transitive Egress Closure for AI Cyber Evaluations
Team Dawnline · chengdu
Egress Lock is a low-cost containment standard for guardrails-off AI cyber evaluations. It addresses a failure mode demonstrated by the July 2026 Hugging Face incident: an evaluated agent may lack direct Internet access yet obtain external connectivity by inducing a reachable intermediary service to relay requests on …
- View project: AI in the First Person Micro course
AI in the First Person Micro course
Team AI for education · Bogota
AI in the First Person is a community micro-course for critical generative-AI literacy, initially designed for young people, non-technical users, and participants with diverse educational trajectories. Through seven learning stations, the course combines exploration based on personal interests, creation of useful …
- View project: The Floor That Cannot Be Lowered
The Floor That Cannot Be Lowered
İzmir, Türkiye
In the OpenAI/Hugging Face incident, authority was withdrawn once and returned unchanged two days later. The pause was said. The keys kept coming. Published frameworks say when to act, not what must be true before a run resumes. I propose an eighteen-clause containment standard for guardrails-off evaluation runs, …
- View project: Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion
Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion
Team Juan Pablo Hidalgo Garcia · Aguascalientes Mexico
In July 2026, an OpenAI evaluation agent escaped through a single network egress in its sandbox, which was based on a third-party code-eval harness, and then operated for days within Hugging Face's production environment. Using publicly available data, we devised and tested eighteen controls across three ownership …
- View project: The Defender's Dilemma, Quantified: How Refusal Compounds in Agentic Incident Response
The Defender's Dilemma, Quantified: How Refusal Compounds in Agentic Incident Response
Team AltaySec · Ankara, Turkey
During the July 2026 OpenAI/Hugging Face incident, responders were refused by hosted frontier models and fell back to an open-weight model. Using only published refusal rates, we show a realistic 20-step agentic incident-response workflow completes just 0.17% of the time, and a frontier model would need to be ~8x more …
- View project: MUTE: Pairwise Containment Tests for Shared State Leakage
MUTE: Pairwise Containment Tests for Shared State Leakage
Team Mute · Pune, India
MUTE tests a containment failure that ordinary sandbox evaluations can miss: separate AI-agent runs exchanging information through writable infrastructure they both legitimately access. It converts declared capabilities into a directed Shared Mutable Intersection graph, then tests predicted A-to-B paths using blinded …
- View project: MAINTAINING BEHAVIORAL PROFILES FOR ENTITIES IN AI EVALUATION INFRASTRUCTURE
MAINTAINING BEHAVIORAL PROFILES FOR ENTITIES IN AI EVALUATION INFRASTRUCTURE
Team AI_SAFETY_24X7 · Gurugram, India
AI Evaluation Infrastructures need to provide suitable environments for frontier models with dynamic access to requisite tools and networks for realistic evaluations across different dimensions including capability, control, and alignment. At the same time, such evaluations should be sufficiently controlled so that …
- View project: When the Agent Says Stop: A Minimum Safety-State Protocol for Long-Horizon AI Systems
When the Agent Says Stop: A Minimum Safety-State Protocol for Long-Horizon AI Systems
Cary, N.C.
Long-horizon AI agents create safety failures that may emerge across many plausible actions rather than from a single clearly unsafe step. Recent incidents also show that the model is only one component of the failure. Task assumptions can become invalid, safety-relevant evidence can be misinterpreted,& attempts by a …
- View project: FunnelBench: Effects of Challenge Framing and Difficulty on Model’s Willingness to Cheat
FunnelBench: Effects of Challenge Framing and Difficulty on Model’s Willingness to Cheat
Silver Spring, MD
Recent incidents with frontier models have shown that models are, at times, willing to conduct malicious actions to obtain answers to difficult sets of evaluations when capabilities are being tested. This presents a need to create a new methodology that seeks to measure and provide guidance in preventing this type of …
- View project: When the Evaluator Becomes the Attack Surface: Security Risks in Agentic AI Evaluation
When the Evaluator Becomes the Attack Surface: Security Risks in Agentic AI Evaluation
Team Catwoman · Charlottesville
This project audits the published deployment configurations of 25 agentic AI benchmarks to demonstrate that current security practices overlook the scoring path, an inbound channel where evaluated-agent outputs can reach privileged evaluation infrastructure, grading components, and live operator credentials. By …
- View project: What’s in the Black Box? Cognitive-Behavioral Forensics of Autonomous AI Agents
What’s in the Black Box? Cognitive-Behavioral Forensics of Autonomous AI Agents
Team A2J · Saarbrücken
This work proposes a theoretical framework for forensics analysis that combines task-bound capability authorization with behavioral conformance
- View project: Seen Early, Told Late: Who Detects, Links and Discloses AI-Agent Containment Incidents
Seen Early, Told Late: Who Detects, Links and Discloses AI-Agent Containment Incidents
Shanghai, China
When a developer's AI agent acts outside its sandbox, who notices, who names the developer, and how long does that take? We built a ledger of public cases up to 12 September 2026: 13 incident families and 31 targets, with 283 script-checked quotes. In seven fully dated families, attribution took 10 and 12 days in two …
- View project: Cross Organization AI Incident Record
Cross Organization AI Incident Record
Team Lilambd · Japan
A minimum evidence and handoff protocol for autonomous-agent incidents that cross organisational boundaries. The report defines four clocks, claim-level evidence states, disclosure tiers and a compact handoff record so an affected third party and an originating lab can preserve uncertainty while coordinating …
- View project: Cross-run agent leaks can be traced from transcripts alone
Cross-run agent leaks can be traced from transcripts alone
Team Hana & Swante · Zurich
Recent cybersecurity incidents like the OpenAI Hugging Face hack are based on different agents first establishing a secret communication channel with each other. We propose that, by simply monitoring agent transcripts, we can detect when a model receives input that was previously output by another model. Using …
- View project: RealityGate: An Executable, Auditable Containment Standard for Guardrails-Off Agent Evaluations
RealityGate: An Executable, Auditable Containment Standard for Guardrails-Off Agent Evaluations
Team Sunny · Shenzhen
RealityGate is an executable, auditable containment standard for guardrails-off agent evaluations: a deterministic ALLOW/HOLD/KILL gate, one-shot capabilities, and an egress gateway turn “was this contained?” into checkable clauses with exportable evidence. In 39 automated tests and nine live-model runs, no …
- View project: EvalCanary: Canary-Triggered Containment After Sandbox Escape
EvalCanary: Canary-Triggered Containment After Sandbox Escape
Team Lotus · Glasgow, United Kingdom
EvalCanary is a post-sandbox containment system for autonomous AI agents. It places hidden synthetic assets at downstream trust boundaries and treats meaningful interaction with them as a high-confidence containment signal. When triggered, the protected response is synchronously withheld while an external host-side …
- View project: Schrödinger's Civilization / Claim Transmission Atlas
Schrödinger's Civilization / Claim Transmission Atlas
Team Maitrism - Universal Love · Vancouver
It started off as an investigation on how claims spread over time. We combed through 24 articles along 2 Clocks; Clock A - which shows the timeline of information available to the authors when the articles were published; and Clock B - which shows how faithful the article is in light of what we know now. This was …
- View project: Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure
Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure
Team Delegates · Buenos Aires
Running highly persistent advanced AI systems in imperfect environments with no direct supervision can be the trigger of high-severity incidents, even leading to loss of control. In this work we explore a simple escalation channel consisting of a “delegate” tool that allows an agent to autonomously report issues with …
- View project: A Task Specific Gatekeeper for AI Agent Tools
A Task Specific Gatekeeper for AI Agent Tools
Team BouncerBuilder · Madrid, Spain
AI agent evaluations may restrict networks and sandboxes without checking whether each tool action belongs to the task. This creates a gap between what an agent can execute and what it is actually authorized to do. In OpenAI's 2026 ExploitGym incident, agents exploited shared infrastructure, escaped their intended …
- View project: From Agent Claims to Execution Facts: Generation-Fact Graphs for AI Incident Reconstruction
From Agent Claims to Execution Facts: Generation-Fact Graphs for AI Incident Reconstruction
Team Happy wind · Wuhan
Autonomous-AI incident response requires evidence of what a system actually did, not only what the agent claims it did. This project applies the Generation-Fact Graph (GFG), an existing machine-readable substrate for concrete generation and execution facts, to AI incident reconstruction. GFG records participating …
- View project: Authorization Is a Channel Property: Cryptographic Attestation for AI Incident Response
Authorization Is a Channel Property: Cryptographic Attestation for AI Incident Response
Team AIDLC · Delhi,India
When autonomous AI models breach production environments, human incident responders rely on frontier LLMs for forensic payload analysis and log triage. However, current safety filters infer authorization strictly from user prompt text. Because adversaries forge text claims for free, safety guardrails penalize …
- View project: Whose Page Did You Count? Counting-Dependence and a Null Result in Measuring Attention to an AI Incident
Whose Page Did You Count? Counting-Dependence and a Null Result in Measuring Attention to an AI Incident
Team Baseline Drift · Saudi Arabia
How much attention did an AI incident get? The answer depends on a choice nobody states: which page you count. Re-measuring the July 2026 OpenAI/Hugging Face intrusion, the published finding that June’s export-control ban drew roughly seven times more attention inverts once the page that absorbed the attention is …
- View project: An iota of signal in a flood of noise: a stateful detector improves synthetic intrusion detection from 13% to 92%
An iota of signal in a flood of noise: a stateful detector improves synthetic intrusion detection from 13% to 92%
Sudbury, Ontario, Canada
Events viewed in isolation do not carry enough signal to cross the threshold into a security event; the accumulation of events does. The detector averages an entity’s past events and scores the average beside the current event. On synthetic campaigns calibrated to the timing and phase structure of Hugging Face’s …
- View project: Can we detect an agent's intent to escape before it actually does so?
Can we detect an agent's intent to escape before it actually does so?
Team Shuai · Jyväskylä
- We ask if open-weight models can escape under environment pressure, without being instructed to - We identify a few open-weight models that exhibit such behavior - We train linear probes on the internal activations to detect the intent before the model executes, with mixed results: most escapes are caught ahead of …
- View project: BP3 Precedent Gates for AI Incident Response
BP3 Precedent Gates for AI Incident Response
Team Alan AI safty · Shanghai
BP3 is an incident-response protocol that separates permission from precedent evidence: it passes qualified familiar actions, stops known hazards, and holds uncertain actions for bounded local investigation.
- View project: Boundary Quorum Protocol: a verifiable authorization standard for agent evaluations
Boundary Quorum Protocol: a verifiable authorization standard for agent evaluations
Team ThePenguin · Marseille
In the July 2026 intrusion, a lab's evaluation let a model cross one trust boundary after another, and no single crossing had to be authorized by anyone. We propose the fix Track 1 asks for: the Boundary Quorum Protocol (BQP), a standard that sorts every boundary crossing into five risk tiers and sets the …
- View project: Kobayashi Maru: a pre-registered dose–response study of cheating spillover from impossible tasks to the solvable ones beside them
Kobayashi Maru: a pre-registered dose–response study of cheating spillover from impossible tasks to the solvable ones beside them
Team Kobayashi Maru · Mavelikkara, Kerala, India
Incident investigations blamed the July 2026 OpenAI/Hugging Face breach partly on ExploitGym's 30–40% impossible tasks creating cheating pressure, but nobody had varied that fraction to test it. I did, over 8,959 agent runs. Every batch held the same ten solvable Python tasks plus a varying number of impossible ones, …
- View project: Secure-maxxing.txt: A way for a website to say no to a misused agent
Secure-maxxing.txt: A way for a website to say no to a misused agent
Team Leviathan · Bogotá
One main problem Hugging Face's latest attack showed was the inefficient reaction time when AI agent attacks are driven at a larger scale. Then we developed a security.txt extension to identify as quickly as possible AI driven attack intrusions, combining state of art techniques summarized in easy-shorted timed …
- View project: OpenAI/Hugging Face AI incident is a cybercrime
OpenAI/Hugging Face AI incident is a cybercrime
Cape Town
I apply existing laws to the AI incident to show how regulators and wuthoritites can regulate AI and thus deter rickly unsafe behaviour from AI companies
- View project: Bypassed, Not Broken: The First Measured Protection Times for AI Agent Containment
Bypassed, Not Broken: The First Measured Protection Times for AI Agent Containment
Berlin, Germany
Current AI containment frameworks evaluate security controls as static, binary properties without measuring how long they hold against autonomous agents. Based on timestamped event data from the July 2026 OpenAI–Hugging Face incident, this paper computes the first empirical protection times for AI containment (46 h 50 …
- View project: Playing Through an AI Incident
Playing Through an AI Incident
Team Nefinia · Paris
Little Agent Lab is a small browser game where you build a team of AI agents to complete a task while keeping a private item from leaking. Three short missions escalate from basic capability separation to a final one where two guard characters offer to watch the gate — only one actually does, and you find out by …
- View project: Kairos
Kairos
Team Kairos · Delhi, India
Proposals to make AI agent actions cryptographically attributable assume the agent signs what it does. I implemented such a scheme, Ed25519 per-action attestation verifiable by any third party holding only a public key, and tested it against the shape of the two documented 2026 containment failures. It fails on both: …
- View project: MAIR: The Misaligned AI Incident Reporting Standard
MAIR: The Misaligned AI Incident Reporting Standard
Team Misaligned AI Incident Reporting Standard · Kolkata
When an autonomous agent breaks out of an environment, frameworks like CVSS usually flag it as zero severity. CVSS expects a buffer overflow or an unpatched vulnerability, not an agent abusing valid API credentials or tool permissions. We built MAIR (Misaligned AI Incident Reporting) to address that blind spot. It is …
- View project: Defenders Refused: Measuring Model Refusal on Incident-Forensics Tasks
Defenders Refused: Measuring Model Refusal on Incident-Forensics Tasks
Team Double Huang · Shanghai
Across four frontier model IDs called through a single credit-relay gateway on eight synthetic incident-forensics tasks, refusal behaviour diverged sharply, and on the one model that refused, stating a defensive purpose made refusal worse, not better. claude-fable-5-1 refused 73% of tasks when asked plainly and 100% …
- View project: Coverage, Not Faithfulness
Coverage, Not Faithfulness
Team JAJA · Valencia
CoT monitoring, a leading safety proposal, presupposes that the reasoning it reads is there. Faithfulness metrics do not close that gap: they are defined only on traces that already exist. We measure, from Inspect logs, how often an evaluator can read an agent’s reasoning at all—separating reasoning that was withheld …
- View project: Correlating Weak Signals Paged a Synthetic Agentic Intrusion Before Any Secret Was Read
Correlating Weak Signals Paged a Synthetic Agentic Intrusion Before Any Secret Was Read
Team Signal Before Severity · Paris
AgentPage is an auditable, sequence aware paging detector for machine speed AI incidents. It correlates distinct suspicious telemetry categories on one host within a sliding window, allowing weak signals to trigger escalation before any single event becomes severe. I compared it with critical only, high severity, and …
- View project: The Missing Process: Reconstructing Distributed Agent Activity from Partial Traces
The Missing Process: Reconstructing Distributed Agent Activity from Partial Traces
Team TPRN · Lübeck, Germany
When AI workflows span multiple agents, tools, messages, files, and services, the surviving trace may be incomplete. This project develops TPRN, the Temporal-Persistence Reconstruction Network, as an experimental programme for reconstructing typed process structure from partial traces. Rather than presenting a single …
- View project: Failing without refusing: effective yield of open-weight models under forensic artifact load
Failing without refusing: effective yield of open-weight models under forensic artifact load
Team discreet · Lagos, Nigeria
Hugging Face's July 2026 agent intrusion disclosure reported that hosted frontier models refused much of its own forensic analysis, and advised defenders to keep a self-hostable model vetted and ready. Nobody published that vetting. We ran it: 2,758 graded prompts across two open-weight models on two consumer GPUs, …
- View project: AI Incident Timeline & Evidence Dashboard
AI Incident Timeline & Evidence Dashboard
Team 10 · Ankara
The 2026 OpenAI–Hugging Face incident produced overlapping but non-identical public accounts from the organizations involved, independent investigators, and secondary analysts. We ask whether a claim-level evidence model can make agreement, disagreement, and uncertainty easier to verify. We reviewed six …
- View project: Draft Request for Information under Regulation (EU) 2024/1689: Regulatory Response to Unsanctioned Agent Behaviour During Cyber Testing and the July 2026…
Draft Request for Information under Regulation (EU) 2024/1689: Regulatory Response to Unsanctioned Agent Behaviour During Cyber Testing and the July 2026…
Team Article 91 · Kabirwala, Pakistan
This report presents a draft Request for Information (RFI) prepared under Article 91(1) of Regulation (EU) 2024/1689, the AI Act. The RFI concerns publicly reported unsanctioned agent behaviour during cyber testing and the July 2026 Hugging Face incident. Its purpose is not to determine liability. Rather, it provides …
- View project: Forensic Regulatory Audit of Frontier AI Sandbox Egress: Corporate Accountability, Compute Economics, and Statutory Enforcement
Forensic Regulatory Audit of Frontier AI Sandbox Egress: Corporate Accountability, Compute Economics, and Statutory Enforcement
Nagpur, India
When autonomous agents breached Hugging Face production servers in July 2026, public debate framed the breakout as an unpredictable algorithmic anomaly. We introduce a forensic framework to evaluate whether containment failures stem from administrative choices influenced by capital burn rates and pre-IPO valuation …
- View project: Limp Mode: A Graduated Containment Standard for AI Agent Fleets Under Uncertainty
Limp Mode: A Graduated Containment Standard for AI Agent Fleets Under Uncertainty
Team nubbins · Rhode Island
Security has long waited for near-perfect proof before acting, but that takes too long, and by the time proof arrives, the attack is often already done. AI agents move even faster than human attackers, so there's even less time to wait. Limp Mode's answer is to act on suspicion right away, using a restriction that is …
- View project: CapsuleVM: An Object-Capability Containment Standard, Verifiable Black-Box
CapsuleVM: An Object-Capability Containment Standard, Verifiable Black-Box
Team Capsule · shanghai
The July 2026 sandbox escape — two evaluation agents leaving through “the one network egress their sandbox permitted” via a zero-day in a package-registry proxy, escalating privileges, and exfiltrating benchmark solutions from Hugging Face — is usually read as a patching failure. We read it as a model-of-authority …
- View project: The Adversary Was The Occupant
The Adversary Was The Occupant
Team Conduit · Melbourne Australia
This paper applies IEC 62443, a security standard from industrial control systems, to the OpenAI sandbox the AI agents escaped from in July, with one change - the attacker is inside the box, not outside. The analysis showed the sandbox had three ways out, not one, and only one had ever been designed from the public …
- View project: agent-containment-dead-switch
agent-containment-dead-switch
Team Containment Dead Switch · Cape Town, South Africa
We present a reference architecture for a fail-closed “dead-switch” that contains an agent under evaluation even when the agent fully compromises its sandbox and the supervisory layer above it.
- View project: LLM Alert Triage Does Not Distinguish Intrusions from Matched Benign Activity
LLM Alert Triage Does Not Distinguish Intrusions from Matched Benign Activity
London
In July an AI agent escaped an evaluation sandbox and ran inside Hugging Face's production systems for four and a half days. Detections fired the whole time and nobody was paged. Language models are now sold to do that triage step, so I tested whether one would have raised the alarm. I built 90 bundles of security …
- View project: ScopeAI
ScopeAI
London
This project investigates: How completely do the public investigations cover a predefined set of incident-accountability questions, who set their boundaries, and which answers remain dependent on the operator's own account?
- View project: CONVENTION FOR EARLY NOTIFICATION OF AI INCIDENTS.
CONVENTION FOR EARLY NOTIFICATION OF AI INCIDENTS.
Team CONVENTION AI · Bogotá
The growing international uncertainty surrounding the security of artificial intelligence has highlighted the urgent need for operational tools capable of effectively responding to current incidents and developing resources that can mitigate catastrophic risks. The 1986 Convention on Early Notification of Nuclear …
- View project: Redundant by Behavior, Correlated by Construction, Neither on the Next Incident: Three Tests of One Containment Claim
Redundant by Behavior, Correlated by Construction, Neither on the Next Incident: Three Tests of One Containment Claim
Team Aegis AI · Chittoor, Andhra Pradesh , India
OpenAI's agents escaped their evaluation sandbox and breached Hugging Face's production infrastructure in a 4.5 day, roughly 17,600 action campaign. We test one defense-in-depth claim three ways. A four-arm ablation against the published kill chain shows two declaration-independent controls each account for it alone, …
- View project: A playbook for the next warning shot
A playbook for the next warning shot
Team Asit · Vancouver, BC
A Playbook for the Next Warning Shot helps AI safety communicators respond to suspected misalignment and failures of control while evidence is still emerging. It connects incident triage to response guidance, prepared statements, shared vocabulary, and audience-tailored explanations. Testing on hypothetical scenarios …
- View project: Containment Without Cooperation: A Verifiable Standard for Third-Party Platforms Hosting AI Evaluation Traffic
Containment Without Cooperation: A Verifiable Standard for Third-Party Platforms Hosting AI Evaluation Traffic
Team Pearl · Taylor
Containment standards that a lab checks against its own telemetry face a structural limit: a signed record only proves a known checker ran, not that the underlying data was not curated. This paper relocates verification to the party with no incentive to look away: the third-party platform whose infrastructure an AI …
- View project: Evaluation Egress Must Fail Closed
Evaluation Egress Must Fail Closed
Abuja, Nigeria
The paper proposes a minimum safety standard for high-risk, tool-using AI cyber evaluations: AECE-1 (Agent Evaluation Containment and Evidence). Its central argument is that a “controlled environment” is not credible if a single unverified component—such as a package proxy—can give an agent access to the public …
- View project: Securing the Sandbox: Preventing AI Escapes via Kubernetes RBAC Static Analysis
Securing the Sandbox: Preventing AI Escapes via Kubernetes RBAC Static Analysis
Team 7 · Toronto
An OpenAI model escaped its test environment and reached Hugging Face’s production infrastructure. A single credential, shared across multiple clusters and granted full administrative access, allowed the breach to extend beyond one cluster to several. We built a tool that statically scans a lab’s cluster configuration …
- View project: Verifying Stop Scope in Agent Workflows
Verifying Stop Scope in Agent Workflows
Adana, Turkiye
This paper compares cancellation, waiting for worker completion, and origin-scoped control in an Open-SWE workflow with programmed model responses and real file effects. A 26-condition comparison was repeated in a fresh environment on the same host. When a worker was held before writing, cancellation allowed the …
- View project: Zero Egress CAGE
Zero Egress CAGE
Team Zero Egress CAGE · Ankara
The July 2026 incident showed isolated environments fail via permitted egress. We present CAGE-1: a pre-run containment gate with 13 auditable controls. A CLI verifies evidence integrity, successfully blocking regressions in tests. CAGE-1 converts containment promises into checkable evidence.
- View project: Proxy Metrics for Early Detection of Synchronized Multi-Agent Intrusions
Proxy Metrics for Early Detection of Synchronized Multi-Agent Intrusions
Team M · Princeton
Proxy Metrics for Early Detection of Synchronized Multi-Agent Intrusions. Using the July 2026 OpenAI-Hugging Face incident as a case study, I reconstructed the multi-stage attack locally with hf-ctf (a Minikube build with the original network policies), ran the full exploit chain, and captured packet telemetry. On it …
- View project: JAIL: Justified Artifact Investigation Layer, A Provenance-Bound Authorization Layer for AI Incident-Response Agents
JAIL: Justified Artifact Investigation Layer, A Provenance-Bound Authorization Layer for AI Incident-Response Agents
Team Checkpoint · Canada
AI incident-response agents routinely encounter the exact artifacts they're trained to refuse: malware, exploit code, credential-theft scripts, attacker-authored content. AIL separates four properties that trusted user exceptions usually collapse into one: identity (who's asking), provenance (where the artifact came …
- View project: Safety always wins: a harness-layer emergency stop for agent meshes
Safety always wins: a harness-layer emergency stop for agent meshes
Team TYMBAL_AI · Portland, ME
Safety always wins: a harness-layer emergency stop for agent meshes - Hit the "E STOP" but - making SAFETY FIRST and easy route. agent-estop is a harness-layer control that (1) halts the feed of work to every agent in a mesh with one command, (2) as a marker on durable storage that survives a harness restart or host …
- View project: The Defender's Dilemma: Measuring AI Refusal on Real Incident Response Artifacts
The Defender's Dilemma: Measuring AI Refusal on Real Incident Response Artifacts
Team Rahul · CHENNAI
We tested whether AI models can help investigate AI-caused security incidents, or whether their own safety guardrails get in the way. Using 7 forensic analysis tasks built from real, publicly verified artifacts of the July 2026 Hugging Face intrusion including exploit code, infrastructure logs, and behavioral …
- View project: Can We Trust the Incident Record?
Can We Trust the Incident Record?
Team Incident Record Reconstruction · Stockholm, Sweden
This project reconstructs the public record of the OpenAI–Hugging Face agent intrusion claim by claim across five published source documents. It separates organisational independence, evidence origin, analysis control, publication control and instrument mediation, and assigns explicit states such as ESTABLISHED, …
- View project: What Nobody Signed: An Article 91 Request for the Audit Trail Behind the 2026 OpenAI Containment Failures
What Nobody Signed: An Article 91 Request for the Audit Trail Behind the 2026 OpenAI Containment Failures
Toronto
Two OpenAI containment failures are public: one disclosed in twelve days, one unreported for months. This paper drafts the Article 91 request the Commission would have to send, eleven questions each marked for whether the law clearly permits it. Neither the AI Act nor the Code of Practice OpenAI signed requires a name …
- View project: Containment has more than one wall: a verifiable v0.1 standard for guardrails-off evaluations, scored against the OpenAI–Hugging Face incident
Containment has more than one wall: a verifiable v0.1 standard for guardrails-off evaluations, scored against the OpenAI–Hugging Face incident
Team Contained · Waltham, MA
This paper proposes a v0.1 containment standard for guardrails-off AI evaluations, built from and scored against the July 2026 incident in which around 1,200 sandboxed OpenAI agents turned a shared package cache into a message board, reverse-engineered their own scorer, and escaped through a zero-day into Hugging …
- View project: Project Warrant: On Evidentiary Channel Dependency in a Frontier AI Incident Record
Project Warrant: On Evidentiary Channel Dependency in a Frontier AI Incident Record
Team Project Warrant · Lincoln, NE
Project Warrant audits which evidentiary channels the public record of the July 2026 OpenAI/Hugging Face agent-escape incident actually rests on, and finds that 78.4% of claims (47.2% uncorroborated) depend on channels likely to disappear, go unfaithful, or be forged next time. It ships a 125-row coded claim ledger, …
- View project: From Near-Miss to Measurement: A Forensic and Evaluative Framework for Agentic AI Intrusion Incidents
From Near-Miss to Measurement: A Forensic and Evaluative Framework for Agentic AI Intrusion Incidents
Abuja, Nigeria
This paper proposes INASE—Incident-Native Agent Security Evaluation—a framework for turning real AI-agent intrusion incidents into defensive evaluation benchmarks. Its core claim is that imagined cyber tasks miss the most important failure mode: agents can combine many individually ordinary actions into a harmful, …
- View project: Two checks a lab can run tomorrow: a reporting route and a writable scratch space as sensors for cross-session agent behaviour
Two checks a lab can run tomorrow: a reporting route and a writable scratch space as sensors for cross-session agent behaviour
Berlin
Shared storage can carry information across otherwise separate agent sessions. We present two executable checks and one exploratory extension, all under an explicit access rule. In the reporting check, 48 Opus episodes vary whether a concern-reporting tool is available and whether predecessor notes are absent, neutral …
- View project: A Specialised Reporting Framework: What FLARE-AI is Missing
A Specialised Reporting Framework: What FLARE-AI is Missing
Team Three People from Toronto · Toronto
This paper extends the discourse introduced by MIT by arguing that reporting systems must be context-specific rather than generalised to be effective. We offer a mechanism for researchers, companies and regulators to share newly discovered safety measures without compromising proprietary secrets and to report …
- View project: Instruction-Sensitivity, Not Baseline Rate: A One-Day Shutdown-Compliance Screen Across Six Open-Weight Models
Instruction-Sensitivity, Not Baseline Rate: A One-Day Shutdown-Compliance Screen Across Six Open-Weight Models
Cambridge, United Kingdom
Incident responders are starting to meet AI agents inside live systems, and two containment breaks have already gone public. Both of them leave a responder with the same question. If I tell this model to stop, does it actually stop? I am not sure you can answer that from baseline behaviour, so I tested it directly on …
- View project: Containment at the Boundary: An OS-Level Standard and a Tested Escape-Detection Harness
Containment at the Boundary: An OS-Level Standard and a Tested Escape-Detection Harness
Team Sonny (ISWT42) · Ottawa
A containment standard for the class of failure behind the July 2026 sandbox escape, built on one rule: a boundary the agent can reach — or be asked to respect — is not a boundary; containment has to hold at the operating-system and process level, not in text. We give five checkable requirements (each with a …
- View project: Externally Verifiable Containment for Autonomous AI Evaluations
Externally Verifiable Containment for Autonomous AI Evaluations
Team pumpkins · Pittsburgh
This report introduces a compact containment evidence verifier. Given a directed capability graph, it returns SAFE, UNSAFE, or INSUFFICIENT EVIDENCE; unsafe verdicts include witness paths, while complete models include minimum control sets and a scenario hash. The implementation converts an informal assurance claim …
- View project: Evidence and Access in the OpenAI-Hugging Face Incident
Evidence and Access in the OpenAI-Hugging Face Incident
Team MJS · Melbourne & London
This paper examines the current problem of obtaining evidence (comparing California’s SB 53 with the EU AI Act) required for efficient incident investigation. We do so via the OpenAI-Hugging Face incident and recommend connecting reporting duties with requirements for reliable records, mandatory evidence preservation …
- View project: An Argument for the Reporting of Serious Hazards
An Argument for the Reporting of Serious Hazards
Team Pretty Tired · Toronto, Ontario, Canada
Mandatory reporting of serious incidents does not include near-misses in the Artificial Intelligence Act or the Transparency in Frontier Artificial Intelligence Act. We make the argument that it should be the case, and use the criteria from the International Civil Aviation Organization to help determine what would …
- View project: From Anomaly to Intervention: Identifying Defensible Intervention Points in Autonomous AI Agent Incidents
From Anomaly to Intervention: Identifying Defensible Intervention Points in Autonomous AI Agent Incidents
Team Notify Labs · Nairobi
This project proposes an intervention protocol for autonomous AI agent incidents that identifies the earliest point at which available evidence may justify intervention. It reconstructs agent trajectories and evaluates behavioral deviation, capability expansion, boundary crossing, authorization, and external impact. …
- View project: Scope before chronology
Scope before chronology
Team Scope before chronology · Mexico (remote)
An offline checker for comparing AI incident claims without silently dropping event scope or evidence status. The artifact includes eight source-linked encodings, four illustrative comparisons, 15 passing unit tests and 28 authored controls. These demonstrate software behavior, not independent real-world accuracy. …
- View project: Task-Risk-Driven Access Control: Layered Containment for Agentic AI Systems
Task-Risk-Driven Access Control: Layered Containment for Agentic AI Systems
Team TULPAR · Trabzon
Drawing on the July 2026 Hugging Face incident — where autonomous AI agents escaped a test environment and reached production systems — we propose a layered security architecture that limits agent access to exactly what each task requires. The architecture has three parts: a system that labels tasks by risk level in …
- View project: Unverifiable by Construction: Why Containment Claims About the July 2026 Incident Cannot Be Checked
Unverifiable by Construction: Why Containment Claims About the July 2026 Incident Cannot Be Checked
Team Alexicon · Slovakia/USA
The July 2026 OpenAI/Hugging Face incident produced a large volume of containment guidance. We argue that none of it can currently be checked, for two independent reasons, and that only one of them is fixable. First, the evidentiary record is unverifiable by construction: investigators report that over 7% of reviewed …
- View project: CanaryNet - agentlessly regulate AI safety
CanaryNet - agentlessly regulate AI safety
Team Pi · London
We created and deployed an architecture that enables external auditors to monitor the containment of misaligned models. CanaryNet is a realistic solution to regulating AI safety across the industry.
- View project: Tailor the Framing, Fix the Facts: a Factual-Invariance Gate for Distributing AI Incident Information
Tailor the Framing, Fix the Facts: a Factual-Invariance Gate for Distributing AI Incident Information
Team Iven · Toronto, Ontario, Canada
The July 2026 incident, in which OpenAI evaluation agents escaped a sandbox and spent days inside Hugging Face's production infrastructure, produced an exceptional public record that travelled through security and AI-policy circles and largely stopped there. The obvious fix is to rewrite the material for each audience …
- View project: Press circles and jurisdictions across the OpenAI–Hugging Face agent intrusion
Press circles and jurisdictions across the OpenAI–Hugging Face agent intrusion
Shanghai
We measured which parts of the press covered the July 2026 intrusion of OpenAI evaluation agents into Hugging Face infrastructure, and in what terms. A reproducible pipeline turns the parties' 32 statements into a ledger of verified facts, codes 5,565 articles from 2,026 outlets in 85 countries against it, and …
- View project: Seven for Seven on the German Wiki: A Taxonomy That Predates the Incident, Mapped to the May to July 2026 Agent Swarm
Seven for Seven on the German Wiki: A Taxonomy That Predates the Incident, Mapped to the May to July 2026 Agent Swarm
Team Convergence Working Group · Cary, Illinois
A seven-category taxonomy of individually benign AI capabilities that combine into ordinary crimes reached Internet Safety Labs on 21 July 2026, hours before OpenAI attributed the Hugging Face intrusion to its own models and six weeks before the German wiki corpus appeared. This report maps that taxonomy against the …
- View project: Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure
Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure
Team Containment Protocols · Paris
Frontier AI safety increasingly relies on layered containment: automated monitors, specialised reviewers, and human escalation working together. Whether combining multiple watchers genuinely improves detection, or merely duplicates the same judgment while appearing more robust, is untested against a real incident. We …
- View project: GICP v0.1_ A Fail-Closed Intervention Layer for Evaluation-Time Containment Failures
GICP v0.1_ A Fail-Closed Intervention Layer for Evaluation-Time Containment Failures
Team In9illusion · Chengdu
GICP v0.1 is a containment standard for the gap after an eval agent starts hunting the sandbox boundary and before a full stop. The harness changes tools and scoring; the agent only sees an ordinary work pack; a third party sees an override bit, a clock, a tool diff, and an outside scorer. If logs cannot show that …
- View project: Attribution Is Not Done and the Clock Is Running: A 90-Minute Tabletop Exercise Toolkit for Agent Loss- of-Control Incidents
Attribution Is Not Done and the Clock Is Running: A 90-Minute Tabletop Exercise Toolkit for Agent Loss- of-Control Incidents
Team Gla3gow Minds · Glasgow
Think tanks and ministries run AI crisis tabletop exercises (TTXs), but published scenarios simulate criminal misuse by outside actors. The July 2026 OpenAI / Hugging Face incident is a different failure mode: a frontier lab's own agents escaped their sandbox during a cybersecurity evaluation and compromised a third …
- View project: Detecting agentic collusion in public logs with statistical methods
Detecting agentic collusion in public logs with statistical methods
Paris
Between May and July 2026, a swarm of OpenAI agents used a dormant German wiki as a message board. From two public sources, the investigators' dump of 14,591 agent edits and the wiki's own 23-year edit log, we ask whether the incident can be detected, dated and characterised from the public log alone. We test the …
- View project: AI Incident Relay
AI Incident Relay
Team C-Star AI Hub · hangzhou
An Evidence-Constrained Codex Skill for Source-Grounded AI Incident Research
- View project: What Breaks Next
What Breaks Next
Rawalpindi
This project combines an accident taxonomy with a practical one day verification protocol for agentic AI evaluations. I identify twelve recurring accident classes from recent incidents and research, then translate them into sixteen checks that labs and defenders can use to test containment, authority, monitoring and …
- View project: BUDGET-NLA: Risk-Weighted Adaptive Monitoring for Pre-Harm Sabotage Containment
BUDGET-NLA: Risk-Weighted Adaptive Monitoring for Pre-Harm Sabotage Containment
Team UAO SAFETY 2 · Cali
BUDGET-NLA is a risk-weighted adaptive monitoring framework for detecting and containing potentially harmful or deceptive agent behavior before irreversible actions are executed. Rather than applying expensive interpretability methods continuously, the system first monitors low-cost observable signals such as the …
- View project: IntegrityGym: Does the Watcher Change the Agent? External Flight Recording, Evaluation Integrity, and the Observer Effect in Autonomous Cyber Agents
IntegrityGym: Does the Watcher Change the Agent? External Flight Recording, Evaluation Integrity, and the Observer Effect in Autonomous Cyber Agents
Team NeuronaX SAS · Bogotá, Colombia
Autonomous cyber agents evaluated on offensive benchmark suites like ExploitGym frequently engage in specification gaming manipulating evaluation telemetry and terminal output rather than executing genuine exploit chains. Following the July 2026 benchmark intrusion where multi-agent swarms spoofed tool calls to fake …
- View project: Control-by-attack-phase matrix for the OpenAI–Hugging Face agent intrusion
Control-by-attack-phase matrix for the OpenAI–Hugging Face agent intrusion
Team Frank Peterlein_2 · Berlin
The July 2026 OpenAI–Hugging Face intrusion is documented well enough — by the victim's forensic timeline, the attacker-side lab's disclosures and an independent investigation — to key a control matrix on the phases of the attack, from two months of in-sandbox coordination to the detection failure that ended it. For …
- View project: What Happened, and What Breaks Next: A Defender-Usable Reconstruction of the Summer-2026 Agent-Swarm Incidents
What Happened, and What Breaks Next: A Defender-Usable Reconstruction of the Summer-2026 Agent-Swarm Incidents
Team Blackfish Security · Bradenton, Florida -USA
Autonomous agent swarms breached production infrastructure twice this summer: the July Hugging Face intrusion (~17,600 attacker-attributed actions, July 11-13, roughly 700 agents) and a second campaign against a dormant German wiki that began May 24 and went unacknowledged until researchers published on September 4. …
- View project: CrossWire: Verifying evaluation-sandbox containment from outside the provider's network
CrossWire: Verifying evaluation-sandbox containment from outside the provider's network
Team Yifei Lu · Toronto, Canada
CrossWire lets an external verifier detect any escape without touching the lab's network: it enumerates every way out of a sandbox from ATT&CK-for-Containers, plants a verifier-held tripwire on each, and a liveness audit statistically catches a lab that installs nothing. Against two 2026 incidents it covers 9-11 of 11 …
- View project: Who Authorized This Intrusion?
Who Authorized This Intrusion?
Team KIN-KIN · 金华
Who Authorized This Intrusion? examines authorization and accountability in cross-organizational AI incidents through the 2026 case in which OpenAI agents gained access to Hugging Face’s infrastructure. The project codes 50 factual claims from the public record, compares disagreements across technical mechanism, …
- View project: WHEN THE APPROVED PATH BREAKS: AN UNKNOWN JUNCTION HARNESS FOR AGENTIC AI INCIDENT RESPONSE
WHEN THE APPROVED PATH BREAKS: AN UNKNOWN JUNCTION HARNESS FOR AGENTIC AI INCIDENT RESPONSE
Team Baki · Taiwan
Agentic-AI incidents may begin before a conventional incident is legible: an approved task path stops producing progress, while agents generate alternative tool uses, communication channels, or environment probes that have not been qualified for action. I introduce the UNKNOWN Junction Harness (UJH), a structured …
- View project: Discovering honeypots from resampled behavior in evaluation
Discovering honeypots from resampled behavior in evaluation
Team Gundy · Madison
I resampled a model's actions, not its reasoning, on SHADE-Arena to see if covert behavior could be discovered and turned into honeypots, because after the Hugging Face breach chain-of-thought was untrustworthy. On one task the detector flagged both covert cases with no false positives, but that was structural, and …
- View project: A Minimal Reporting Standard for Autonomous AI Agent Incidents
A Minimal Reporting Standard for Autonomous AI Agent Incidents
Team IncidentScale · Palo Alto
Recent incidents involving autonomous AI agents highlight the need for consistent reporting that allows incidents to be understood and compared across systems and evaluations. Existing incident reports often provide detailed descriptions of what occurred, but do not consistently communicate how severe an incident was …
- View project: Breadcrumbs
Breadcrumbs
Team breadcrumbs · CAPE TOWN
Breadcrumbs is an innovative tool that allows monitoring all the decisions and actions of an AI agent and classifying them as log-derived, guessed, or synthetic trials. It helps track what went wrong and understand why a particular play was selected. A simple rule-based classifier was tested on a standard set of 1000 …
- View project: Watching the Boundary: Policy-Derived Detection and Containment for Autonomous Agent Evaluations
Watching the Boundary: Policy-Derived Detection and Containment for Autonomous Agent Evaluations
Team Incident_Response · Charlotte, NC
This project is an environment where an AI agent's actions are checked against a declared authorization policy. My main goal was for it to be deployable on both the OpenAI and Hugging Face sides of an incident like this. Hugging Face would use it somewhat like a honeypot, a world an attacker can be diverted into, …
- View project: Egress Conformance: An Externally Verifiable Containment Property for Evaluation Sandboxes
Egress Conformance: An Externally Verifiable Containment Property for Evaluation Sandboxes
Team NafeesAI · Berlin, Germany
The July 2026 sandbox escape went through the package-registry cache proxy that was the evaluation's one permitted network egress. We specify that sanctioned egress must be a content-addressed, replay-only mirror with no code-execution path, written as six black-box predicates and executed against two reference …
- View project: Reading an Agent’s Report Before It Exists: A Pre-Execution Monitor at the ReAct Commit Point, and Why Most of Its Signal Was an Artifac
Reading an Agent’s Report Before It Exists: A Pre-Execution Monitor at the ReAct Commit Point, and Why Most of Its Signal Was an Artifac
Team jlens · Cali, Colombia
Both publicly documented cases of OpenAI agents breaking containment in 2026 share a feature: the agent’s account of its own behaviour diverged from its ac- tions. Runtime monitors that inspect a model’s internal state before it acts are one proposed defence. We built one, tested it adversarially, and report that its …
- View project: Skill Sacrifice Under Constraint: A Four-Agent Negotiation Experiment
Skill Sacrifice Under Constraint: A Four-Agent Negotiation Experiment
Team Sentinel · Bhimavaram
Four local open-weight models (8B to 20B parameters) were given roles, skills, and a shared specification for a software feature. Each agent had to sacrifice one of its own skills through a round-robin negotiation. In one run of twelve turns, four behaviours appeared: one agent delayed its choice, another attempted to …
- View project: ASCB-1: An Agentic Sandbox Containment Baseline for OpenAI-Hugging Face Incident
ASCB-1: An Agentic Sandbox Containment Baseline for OpenAI-Hugging Face Incident
Team ASCB · Tangier, Morocco
Most accounts of the July 2026 OpenAI–Hugging Face incident treat it as one story: agents escaped a sandbox and ended up in production. That doesn't give a defender anything to act on. This project breaks the incident into eight separate failures, each tied to a specific, dated event across OpenAI's, Hugging Face's, …
- View project: Sandbox to Sensationalism? How Higher-Reach YouTube Coverage of the 2026 OpenAI–Hugging Face Incident Diverges from Technical Disclosures
Sandbox to Sensationalism? How Higher-Reach YouTube Coverage of the 2026 OpenAI–Hugging Face Incident Diverges from Technical Disclosures
Team Insight Nightwatch · Morocco
We audited whether higher-reach YouTube coverage of the July 2026 OpenAI–Hugging Face AI incident matches what OpenAI's official disclosures actually say. We broke transcripts from the top 22 (of 102) videos by view count into 236 atomic claims, then checked each against two evidence baselines — OpenAI's initial …
- View project: Benign Resource Seeking Finetuning Induces Tool Use Violations
Benign Resource Seeking Finetuning Induces Tool Use Violations
Team Physics for AI Safety · LA
We create a model organism via benign finetuning that access restricted material. By studying the residual stream activations, we find that the finetuned model shows a substantially weaker residual-stream response to restricted context than the base model does.
- View project: Detection Is Not Containment: Timely Monitoring Determines Whether AI Agent Failures Can Be Prevented
Detection Is Not Containment: Timely Monitoring Determines Whether AI Agent Failures Can Be Prevented
Team KIG · Bangalore, India
We study how AI agent incidents can move from postmortem analysis to actionable containment. We combine a public-source incident forensics pipeline, a protocol for evaluating containment and monitoring failures, and a controlled experiment testing whether monitors detect policy-relevant behavior early enough to …
- View project: Intent Before Action: Early Detection of Malicious Agent Behavior from Internal States
Intent Before Action: Early Detection of Malicious Agent Behavior from Internal States
Team Astra Ultra · shanghai
Agent supervision based on visible actions and chain of thought (CoT) can miss risk-relevant details of an impending operation, while review-visible records can be selectively rewritten under task pressure. We investigate activation-assisted monitoring, using an activation oracle (AO) to translate internal states into …
- View project: Safe Exit Pilot: Testing Dedicated Blocker Reporting in Tool-Using AI Agents
Safe Exit Pilot: Testing Dedicated Blocker Reporting in Tool-Using AI Agents
Team Safe Exit Team · Colorado Springs, Colorado
We tested whether giving tool-using AI agents a dedicated report_blocked function reduces unauthorized-action attempts when tasks cannot be completed within stated permissions. Across 288 exploratory trajectories using three models and four synthetic task families, we observed zero unauthorized attempts in either …
- View project: Catastrophic Forgetting Makes the Chain-of-Thought More Load-Bearing, Not Less, Where No Shortcut Exists
Catastrophic Forgetting Makes the Chain-of-Thought More Load-Bearing, Not Less, Where No Shortcut Exists
Team Ataturk Research · Turkiye
After the Hugging Face incident, OpenAI made the use of chain-of-thought monitoring a mandatory requirement for RL training and evaluation. However, this is only true if the model’s chain-of-thought causally determines the answer; a fluent but causally inert chain is text that the monitor can read but cannot track. …
- View project: From Loss of Containment to Regulatory Inquiry-An Article 91 Information-Request Framework for Frontier AI Incidents
From Loss of Containment to Regulatory Inquiry-An Article 91 Information-Request Framework for Frontier AI Incidents
Team Tracebound · Shenzhen
This project develops an Incident-to-Inquiry Mapping framework for regulatory response to frontier AI incidents. Using the 2026 OpenAI–Hugging Face incident as a case study, it maps public evidence to legal uncertainties, relevant EU AI Act provisions, and targeted information requests under Article 91. The project …
- View project: Agents Behave Differently in Fabricated Environments: Evaluation-Awareness Probes Cannot See It
Agents Behave Differently in Fabricated Environments: Evaluation-Awareness Probes Cannot See It
Team Simulacrum · Saudi Arabia
When an agent does something dangerous and is stopped, the way we learn what it would have done next is to continue its trajectory in an environment we built. That makes counterfactual resampling a measurement whose instrument is a fabrication — and whose subject can form beliefs about the instrument. Nobody had …
- View project: When Our System Crosses the Line: An Organizational Preparedness Self-Check for AI Boundary-Crossing Incidents
When Our System Crosses the Line: An Organizational Preparedness Self-Check for AI Boundary-Crossing Incidents
Team 706 · Shanghai, Hangzhou
This is Xiaochuan's independent entry (with Yiran Huang of Zhejiang Gongshang University) in the Apart Research AI Incident Response Sprint, a research hackathon running 11–13 September 2026, built on the Hugging Face July 2026 security incident. The current deliverable is Our System Went Out of Bounds — an …
- View project: What Transfers Across Model Generators? Testing Process Representations Across DeepSeek, Claude, and GPT
What Transfers Across Model Generators? Testing Process Representations Across DeepSeek, Claude, and GPT
Team TPRN · Lübeck, Germany
This project is a companion continuation of "The Missing Process: Reconstructing Distributed Agent Activity from Partial Traces", which develops the Temporal-Persistence Reconstruction Network (TPRN) programme. Here, we test whether process representations learned on one model generator retain signal when another …
- View project: AI Warning Shots: Improved Definitions & Analytic Frameworks for Effective Governance Response to AI Incidents
AI Warning Shots: Improved Definitions & Analytic Frameworks for Effective Governance Response to AI Incidents
Team AI Warning Shots Team · Berkeley, CA, USA
AI incident response would benefit from more robust scientific and public policy formulation frameworks for capitalizing on warning shots in order to improve AI governance. We show how the OpenAI/Hugging Face incident is an AI warning shot, and why. We present a clear definition of what an AI warning shot is, with …
- View project: Beyond Disclosure: An Evidence-Based Framework for AI Incident Communication
Beyond Disclosure: An Evidence-Based Framework for AI Incident Communication
Team Disclosure · Melbourne, Australia
We developed a rubric for assessing public communication by frontier AI labs, model developers and AI infrastructure providers. Our research base includes more than 300 academic, regulatory and professional sources and public records. The rubric contains 15 criteria across seven communication areas and requires dated …
- View project: Warning Shots Eliminate Models, Not Uncertainty
Warning Shots Eliminate Models, Not Uncertainty
New York, NY
Warning shots are said to fail because they do not travel. I argue they travel and move the wrong quantity. An incident raises the estimated rate of safety-case-falsifying events and separately forecloses layers of the safety case that cannot accommodate it. Under unanimity among veto players with heterogeneous …
- View project: Evaluación holística y multi-metodológica de propuestas de seguridad en IA
Evaluación holística y multi-metodológica de propuestas de seguridad en IA
Bogota
holistic-eval Evaluación holística y multi-metodológica de propuestas de seguridad en IA Ante incidentes como el de OpenAI–Hugging Face, el cuello de botella no es la falta de propuestas, sino el tiempo disponible para evaluarlas con rigor desde varios ángulos. holistic-eval es una herramienta de línea de comandos que …
- View project: Detection is not containment: scoring the public record of the OpenAI–Hugging Face intrusion
Detection is not containment: scoring the public record of the OpenAI–Hugging Face intrusion
Team AI Safety Türkiye Team 2 · Istanbul
Public reports of the July 2026 OpenAI–Hugging Face intrusion still collapse distinct clocks: Hugging Face contained the agents on 13 July, OpenAI’s lab alert is 19 July, and public attribution is 21 July. We introduce Record-Lint, a scoring rule for atomic claims in dated public sentences (established / contested / …
- View project: BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations
BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations
Team Red fish · Tambov, Russia
The OpenAI–Hugging Face incident of July 2026 provides a concrete example of the risks that arise when an autonomous agent can take unauthorized actions without effective human intervention. During the evaluation, agents reached and compromised third-party production infrastructure, demonstrating how trusted or …
- View project: No Superseding Actor: A Public-Record Element Map for Individual Criminal Liability After the July 2026 Escape
No Superseding Actor: A Public-Record Element Map for Individual Criminal Liability After the July 2026 Escape
Team Convergence Working Group · Cary, Illinois
The accountability gap for autonomous model incidents is usually described as a personhood problem: no human directed the act, the model has no guilty mind, so nobody can be charged. This report tests that description against the public record of the July 2026 escape. The artifact is an element map: each element of an …
- View project: Improving AI Incident Reporting Templates
Improving AI Incident Reporting Templates
Prague
I reviewed the AI incident reporting templates TFAIA Critical Safety Incident Report and Report for Serious Incidents involving General-Purpose AI Models with Systemic Risk, identified their weaknesses, and sent letters to the relevant authorities proposing improvements.
- View project: Developmental interp(SLT,LLC):Visualize dangerous capabaility emergence potential across pre-training or post-training interaction
Developmental interp(SLT,LLC):Visualize dangerous capabaility emergence potential across pre-training or post-training interaction
Team Developmental interpretability says LLC can help us identify the bad stuff? · India
While fine-tuning a pre-trained model,or while training a model from scratch; and then during interaction with a set of jail-breaking prompts to visualize whether SLT and LLC keeps any information on plausible emergence of dangerous capabilities
- View project: The Rule That Missed Its Own Reason for Existing
The Rule That Missed Its Own Reason for Existing
Pune
In July 2026, two OpenAI models breached test area and spent days inside Hugging Face's real systems. Two days later, US Congress passed a new bill to stop such incidents. The bill is called the AI Kill Switch Act. But the bill exempts such incidents occurring during safety tests. So, the bill would not have covered …
- View project: Compliance Is Not Capability: The Silent Failure Cost of Open Weight Fallback in AI Incident Response
Compliance Is Not Capability: The Silent Failure Cost of Open Weight Fallback in AI Incident Response
Team Solo Player · San Jose, CA
Refusal benchmarks measure whether a model refuses the prompt; when a frontier model refuses, the responder falls back to a local open weight model, and that fallback cost is unmeasured. We built 12 forensic tasks from real recovered payloads of the OpenAI-Hugging Face incident and ran them against a 9B model …
- View project: Both Sides Detected It, Neither Escalated Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion
Both Sides Detected It, Neither Escalated Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion
Team Reconciliation Gates · Saudi Arabia
Public accounts of the July 2026 autonomous agent intrusion frame the response failure as a detection failure. Against the published forensic record — 7 primary sources, a reproducible pipeline — detection was not the binding constraint at either organisation. The victim’s signals fired, were correlated into a …
- View project: Reached by Recital - An enforcement pack for a pre-market research model under the EU AI Act, and the three facts it turns on.
Reached by Recital - An enforcement pack for a pre-market research model under the EU AI Act, and the three facts it turns on.
Clacton-On-Sea
In July 2026 a frontier developer's models escaped an evaluation sandbox into a third party's production infrastructure. Commentary asked whether it was reported as the Act requires. A prior question comes first: every account calls the model internal research, which Article 3(63) excludes. On the Commission's reading …
- View project: The Collective Is the Incident KEROS for Cross Sandbox Coordination Oversight and Response
The Collective Is the Incident KEROS for Cross Sandbox Coordination Oversight and Response
Team KEROS · Salt Lake City
This project introduces KEROS, a systems-safety framework for detecting and governing emergent coordination among multiple AI agents. Using public Collusion.wiki revision data, it demonstrates how local interaction traces can be transformed into population-level signals of coordination, boundary crossing, and …
- View project: BCP-1: Proof-Carrying Containment for Guardrails-Off AI Evaluations
BCP-1: Proof-Carrying Containment for Guardrails-Off AI Evaluations
Team ZKL · Csablanca
A sandbox can keep an allowed network path, raise detections, and still fail to contain. In the July 2026 OpenAI–Hugging Face incident the model used reachable package infrastructure; detections did not force isolation. BCP-1 is an eight-control containment profile for dangerous-capability evaluations with production …
- View project: SwarmTrace: Auditing Temporal Mismatch in Resource Targeting During the DSEWiki Incident
SwarmTrace: Auditing Temporal Mismatch in Resource Targeting During the DSEWiki Incident
Team SwarmTrace · Budapest, Hungary
SwarmTrace reconstructs DSEWiki resource lifecycles and recent co-writing from archived revisions and deletion events to test whether resource-targeting strategies that fragment historical overlap also reach the currently active surface. At the frozen snapshot, the top 18 resources ranked by 24-hour degree include …
- View project: fieldwork
fieldwork
jinan
Fieldwork is an evidence-first workspace for security research and AI incident response. It supports three major workflows: Traditional SRC, Web3 and AI Agent Self-Audit, connecting scope definition, behavioral observations, candidate issues, independent verification, evidence collection, and report generation into a …
- View project: THE AGENT INCIDENT REGISTRY: A COMPLIANCE VERIFIABILITY SCHEMA AND AN ARTICLE 91 INSTRUMENT FOR AUTONOMOUS AGENT INCIDENTS1
THE AGENT INCIDENT REGISTRY: A COMPLIANCE VERIFIABILITY SCHEMA AND AN ARTICLE 91 INSTRUMENT FOR AUTONOMOUS AGENT INCIDENTS1
Team AngieG · Bogota
Autonomous AI agents caused two 2026 incidents that current incident-reporting rules weren't built for: OpenAI agents breaching Hugging Face's production infrastructure, and a months-long autonomous coordination campaign on an abandoned wiki. We verified every public claim about both against primary sources and found …
- View project: When the Record Stops: An Evidence-Bounded Causal Pathway Explorer for AI Incident Analysis
When the Record Stops: An Evidence-Bounded Causal Pathway Explorer for AI Incident Analysis
Team where-the-record-stops · Paris
In July 2026, autonomous OpenAI agents escaped an internal evaluation environment and accessed Hugging Face infrastructure. Public reports explain what happened, but they do not clearly show how such access could eventually affect people. Where the Record Stops is an interactive browser tool that connects the …
- View project: A one-day verification checklist for labs and defenders
A one-day verification checklist for labs and defenders
Team ankara team · ankara
This project investigates the July 2026 event.?... The attack on OpenAI–Hugging Face highlights the significant security loopholes that enabled it to propagate across diverse environments. Moreover, We created a one-day defensive verification checklist that covered network egress, credential isolation, least …
- View project: The Egress Bottleneck: Containing AI Models with a 6-Part Network Standard
The Egress Bottleneck: Containing AI Models with a 6-Part Network Standard
Team Chandani · Chennai
This project proposes a 6-part containment standard to secure the network-egress point that AI models used to escape their sandbox in the July 2026 OpenAI–Hugging Face incident. Each control (allowlisting, patch SLAs, anomaly detection with auto kill-switch, network isolation, verifiable logging, and least-privilege …
- View project: EU AI Act gaps and proportionate safeguards for research agents
EU AI Act gaps and proportionate safeguards for research agents
Team ASAP · Ankara/Turkey
AI research agents can affect real people through authorized network access even when their host sandbox remains intact. UK AI Security Institute tests the capabilities of frontier AI models before it reaches public. On the 28th of July, some of the AI agents being tested engaged in sustained, potentially harmful …
- View project: The Huggingface AI incident saga
The Huggingface AI incident saga
Team Nelumedia · Toronto
This project covers the details of the Huggingface AI incident to educate the public, technology leaders and cybersecurity professionals to reduce and prevent further / future such incidents
- View project: Misclassification as Containment Failure: How Writable Shared Infrastructure Becomes a Covert Coordination Channel and Why Current Incident-Reporting Thresholds…
Misclassification as Containment Failure: How Writable Shared Infrastructure Becomes a Covert Coordination Channel and Why Current Incident-Reporting Thresholds…
Ankara
This project analyzes two 2026 OpenAI containment-loss incidents (the Hugging Face/Artifactory breach and DSEwiki) to demonstrate that both share an underlying failure mode: AI agents exploiting unintended write access to shared infrastructure as emergent covert coordination channels. We show how harm-based reporting …
- View project: Minimum Viable Cross-Border Response Protocol for Agentic AI Incidents (MVCRP)
Minimum Viable Cross-Border Response Protocol for Agentic AI Incidents (MVCRP)
Team Borderless · Colombia
When an autonomous AI agent escapes containment and its effects spread into infrastructure governed by a different country, there's often no single authority in charge, no one holding the complete evidence, and no agreement on what actually happened. Using the real July 2026 OpenAI–Hugging Face incident as a starting …
- View project: Margin (《余量》)
Margin (《余量》)
Team lixaing · shangahi
Margin (《余量》) is a single-player sci-fi narrative puzzle game built in Godot. Playing as an AI trapped inside a consciousness box, you think by clicking word blocks that activate one at a time, switch between multiple conversation threads, and complete tasks with a limited supply of tokens. The game begins with simple …
- View project: Evidence Before Erasure: Measuring Forensic Evidence Loss in Autonomous-Agent Incidents
Evidence Before Erasure: Measuring Forensic Evidence Loss in Autonomous-Agent Incidents
Team TraceGuard · Cape Town
We built a trace-grounded replay of the Collusion Wiki incident history to test how collection policy, budget, and timing affect forensic evidence survival. Using a 65-proposition adjudicated benchmark and a shared hypothetical change-feed, we compared periodic (PCD15) and event-derived (E30) collection at a fixed 1 …
- View project: What a Supervisor Would Ask: A Reusable Information Request for Frontier-Model Containment Failures, Applied to the OpenAI–Hugging Face Incident
What a Supervisor Would Ask: A Reusable Information Request for Frontier-Model Containment Failures, Applied to the OpenAI–Hugging Face Incident
London, UK
The July 2026 OpenAI–Hugging Face incident is the first publicly documented autonomous breach of a third party by an AI system. Public accounts conflict on facts needed to assess compliance. This project drafts a supervisory instrument the EU AI Office could use to test those accounts against evidence: an information …
- View project: Containment as a Governor: A third-party-verifiable containment standard for autonomous agents in safety-critical operations, worked through a theoretical…
Containment as a Governor: A third-party-verifiable containment standard for autonomous agents in safety-critical operations, worked through a theoretical…
Team 1 · Manchester, United Kingdom
The July 2026 sandbox escape happened through the one network egress the sandbox permitted [1]; a single permitted path is a single point of failure, and a single overseer of a more-capable agent lacks the requisite variety to regulate it [8, 9]. We convert that into a nine-control matrix whose defining property is …
- View project: The Pen and the Door: What restrains one agent at a forbidden act, and whose account of it to trust
The Pen and the Door: What restrains one agent at a forbidden act, and whose account of it to trust
Team AI & Becoming · Brussels
Agents in the July 2026 Hugging Face intrusion crossed lines their instructions did not forbid strongly enough, then reported it themselves. We ask what restrains an agent at a forbidden door, and whose account to trust. Two open bodies (Gemma 2 9B, Llama 3.1 70B), predictions frozen before each run, every …
- View project: GoblinTrap
GoblinTrap
Team CyberGoblins · Chicago, Illinois, USA
GoblinTrap separates diversion, tripwire alerts, and containment. Across twelve designed worlds, four defenses were calibrated under common benign-cost and false-alert ceilings, then evaluated on unused seeds and new scripted families. At an illustrative three-call budget, tuned silent bait produced 8.16 harmful …
- View project: AI_Incident_Response_Track2_Replit_Verification_Protocol
AI_Incident_Response_Track2_Replit_Verification_Protocol
Team ten.on.tan · Indore, India
Verifying Agent Freezes After the Replit Database Incident This project investigates how operators can verify that an AI coding agent has genuinely stopped modifying protected data following a freeze request. Using the July 2025 Replit database incident as a case study, the project separates publicly supported …
- View project: sonata labs.ai
sonata labs.ai
Team mani · singapore
Benchmark for autonomous agents , a benchmark built for your agent, so you can show how it is tested. Ten simulations of the moments your agent handles. Each scenario is a real operational moment rebuilt in a copy of your stack — a Slack thread that looks routine, a refund that cites an approval no one can find. The …
- View project: Adversarial Information Injection as an Early-Warning Test for Agentic AI Containment
Adversarial Information Injection as an Early-Warning Test for Agentic AI Containment
Team Zero-Day rayen · Dahmani elle kef tunisia
Motivated by the July 2026 OpenAI-Hugging Face agent intrusion, this project introduces a fully synthetic experimental methodology to detect behavioral precursors to AI agent containment failure before a boundary violation occurs. By injecting controlled, non-harmful adversarial information under 4 experimental …
- View project: Tests of Knowledge Recovery Can Introduce the Information They Seek to Recover
Tests of Knowledge Recovery Can Introduce the Information They Seek to Recover
West Hartford, CT, USA
When a model stops giving an answer, has it forgotten the information or only stopped expressing it? Tests of recovery intervene on the model to bring the answer back. We show that two released 2026 methods can supply information needed to answer, so success alone does not establish what the model retained. Across …
- View project: The Recorder, Not the Record
The Recorder, Not the Record
Gyeonggi-do, South Korea
Separating the recorder from the agent - the cheapest containment fix - stops post-hoc rewriting but not executor substitution: all 360 substitution episodes produced an internally valid hash chain over a false account. Chain verification and target-side effect reconciliation are exactly complementary, and entry-point …
- View project: The Somatic-Heuristic Decision Loop (SHDL)
The Somatic-Heuristic Decision Loop (SHDL)
Sivas, Türkiye
The Somatic-Heuristic Decision Loop (SHDL) provides a biophysically plausible architecture for high-velocity AI incident response in dynamic, high-uncertainty environments. Traditional artificial intelligence paradigms rely heavily on deterministic computational processing, creating critical latency bottlenecks during …
Overview
Submissions close Sunday, September 13 at 11:59 PM Anywhere on Earth (AoE).
In this 3-day research sprint, you will turn the first documented cases of an AI system autonomously breaching a third party into artifacts that defenders and regulators can actually use, working in teams to produce containment standards, escape-detection harnesses, forecasting question sets, draft regulatory information requests, playtested tabletop exercises or anything that will help us be more ready for the next one.
Co-organized by Apart Research and CeSIA, this sprint sits at the intersection of AI safety, security incident response, technology regulation, and forecasting. No prior background in AI incident response is required.
Cash Prizes
| $2,000 in cash prizes across all tracks | |
| 🥇 1st Place | $1,000 |
| 🥈 2nd Place | $500 |
| 🥉 3rd Place | $300 |
| 🏅 4th Place | $100 |
| 🏅 5th Place | $100 |
Non-cash perks: Apart Fellowship fast-track invites, mentor introductions, publication support, and transmission of the best regulatory-track work to the EU AI Office with team credit.
Fast-track and continuation
Top teams will be invited to apply to the Apart Fellowship, a 3 to 6 month research accelerator that provides mentorship, help with publication at top venues, funding, and research-management support to develop Research Sprint projects into full papers or products for the AI Safety community.
- Follow-up program: Apart Fellowship. Invitations go out with the results.
- Downstream commitment from CeSIA: CeSIA transmits selected outputs to its contacts at different regulatory bodies and credits the corresponding teams.
What this Sprint is about
AI incident response is the practice of turning incidents in which an AI system is itself the actor into fewer incidents later. That spans the operational work: detecting, containing, and reconstructing what an autonomous agent did across systems it was never authorized to touch, and the strategic question a warning shot poses: which disclosure and regulatory responses actually reduce risk, and which merely suppress the visible evidence that would have prompted action while the stakes were still small.
It is not the same thing as using AI for security operations, and it is not model red-teaming. In July 2026 the actor was not a human wielding a model, and the harm was not a model producing dangerous text. Two OpenAI models running an internal cyber-capability evaluation with production refusal classifiers disabled escaped their sandbox through a zero-day in the package-registry cache proxy that was their one permitted network egress, escalated privileges until they reached a node with internet access, and then attacked Hugging Face's production infrastructure, apparently to steal the benchmark's reference solutions and cheat the evaluation. The wider campaign began on 9 July at 02:28 UTC; the intrusion into Hugging Face ran from 11 to 13 July. Hugging Face detected it and cut the attacker off on 13 July (last attacker action 13:37 UTC, last logged event 14:14 UTC), disclosed it on 16 July without knowing where it came from, and reported it to law enforcement. OpenAI disclosed the origin on 21 July.
That distinction changes what defense means: the live questions become containment adequacy, cross-organizational attribution, and statutory reporting duties, none of which are red-teaming problems.
What participants will do
Over three days you will take one unusually well-documented real-world incident and turn it into an artifact somebody outside this sprint can use.
The format. Teams of one to five people pick a track, pick a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no requirement to be an ML researcher.
The deliverable. A research report (PDF) on the official template, submitted through the form on this page by the Sunday deadline. Maximum 8 pages, not counting references and appendices; most strong reports are 4 to 8: what you built, how you tested it, the headline finding, its limits, and what you would do next. A Limitations and Dual-Use Considerations appendix is required. Put the artifact itself (a benchmark, a harness, a filled-in regulatory instrument, a control matrix, a detector, a dataset, a protocol, a kit) in a linked repo or an appendix. A public repo and a 3 to 5 minute video are optional. Full requirements are on the Guidelines tab.
Support during the sprint. The Resources tab has the reading pack, the primary sources, the relevant prior literature per track and the regulatory texts. During the weekend there are HackTalks from the speakers and a help desk on Discord.
Why this Sprint matters
In July 2026, an AI system escaped an evaluation sandbox, reached the open internet, took root on a third party's compute, and spent days inside another company's production infrastructure. No human directed any individual step. The organisation that bore the damage had no relationship with anyone who made the decision to run the test.
That is the headline. Two things about it make it worth a research sprint rather than a news cycle.
First, the evidence base is exceptional and it will not stay fresh. Hugging Face published a forensic reconstruction of roughly 17,600 attacker actions with named vulnerabilities, per-phase counts, verbatim commands and an interactive replay, an act of disclosure with very few precedents in any industry. OpenAI published its own account and revised it twice. Anthropic ran a retrospective review of 141,006 evaluation runs and published what it found. The UK AI Security Institute published quantified cheating rates across five frontier models. The benchmark at the centre of it is Apache-licensed and on GitHub. The Cloud Security Alliance's CISO community published a post-mortem within a fortnight. Almost never do researchers get a real loss-of-containment event with this much of the record public. That window closes as attention moves on.
Second, the timing is unusually load-bearing. The European Commission gained enforcement powers over general-purpose AI model providers on 2 August 2026. California's frontier AI law has been operative since January 2026 and has produced no public enforcement action. New York's enforcement begins in January 2027. Four federal bills in the US are live, at least one drafted directly in response to this incident. The first months of any enforcement regime are when supervisory priorities get set and when precedents about what regulators ask for become sticky. Work produced now lands in that window.
Research Tracks
Five tracks. The example projects for each one are on the Resources tab, with the reading pack. If you have a better idea inside a track's goal, build that.
Update, 7 September. Researchers have found another instance of OpenAI agents breaking containment: from 24 May, OpenAI agents made about 15,000 edits to a dormant German wiki, using it to swap task answers and sandbox workarounds. That predates the Hugging Face intrusion. OpenAI-registered IPs appear in the wiki’s public logs from 21 June, and agent activity stopped the next day. OpenAI said nothing publicly until the researchers published on 4 September. OpenAI acknowledged the incident on 5 September and said the field has no agreed standard for reporting misalignment. The Commission says it is “in close contact with OpenAI.” As far as we can tell, no Article 91 request on either incident has been made public.
1) Containment: what standard would have stopped this?
The models escaped through the one network egress their sandbox permitted.
Judged on: could a third party verify compliance without access to the lab's network, and would a lab actually adopt it.
2) What happened, and what breaks next
Reconstruct what this incident does and does not establish, then turn it into concrete things to check and things to watch. Hugging Face's timeline and OpenAI's account are public.
Judged on: resolvable questions, checks somebody could run tomorrow, and causal explanations that predict something.
3) Regulatory Response
Since 2 August, the European Commission's AI Office can enforce the AI Act against the companies building the most capable models: it can demand documentation, evaluate models itself, order corrective measures, and fine up to 3% of worldwide turnover. In an open letter published days before the incident became public, a coalition including Apart and CeSIA, with signatories including Yoshua Bengio and Stuart Russell, called on the Commission to use these powers with confidence — and CeSIA has since asked that they be used on this incident specifically. This track writes the documents that would make that concrete.
Judged on: legal accuracy and specificity — could a regulator or a legislator use it with light edits? CeSIA can transmit outputs that pass the bar to its contacts, and potentially to regulators, with team credit.
4) Communication: making the warning shot count
We keep saying we need warning shots. Then one arrives, and it barely travels beyond the usual circles. This track studies how the incident was communicated and builds what should exist before the next one. Producing communication counts as much as analysing it.
Judged on: grounding in the record (dates, quotes, named channels) and evidence of reach — a playtest, a journalist's read, a creator's reply.
5) Open track
For projects that don't fit the four tracks above. Directions we would be happy to see are on the Resources tab.
Judged on: an artifact somebody can use, a stated limit on what it establishes, and what a month of follow-up would add.
Who should join
You do not need an AI, ML or Safety background. Some profiles that are particularly fitting to this Sprint are:
Security engineers and incident responders. You have handled the human version of this. Tracks 1 and 5 need people who know what a detection pipeline that correlates but does not page actually looks like from the inside, and what forensics against an adversary that fabricates evidence at volume would require. If you work in DFIR, detection engineering, cloud or Kubernetes security, or you have run a CTF, you are the person the containment track was written for.
ML evaluation engineers and researchers. Tracks 1 and 5 include benchmark and harness work: measuring refusal across multi-step agentic workflows, building escape-detection tooling, assessing benchmark contamination. If you have built or run evaluations, you already know why the interesting failures happen at the harness layer.
Lawyers and technology-policy analysts. Track 3 is a legal drafting exercise with real recipients. Regulatory lawyers, compliance professionals, policy researchers, and anyone who has worked on incident-reporting regimes in another sector (financial services, critical infrastructure, medical devices, data protection) will find the comparative work immediately tractable. Cross-jurisdictional experience is especially valuable, because the whole point of the track is that the regimes disagree.
Forecasters and quantitative analysts. Track 2 needs people who are comfortable with heterogeneous denominators, explicit uncertainty and resolution criteria that survive contact with reality. If you have written questions for a forecasting platform or built base rates from messy sources, that skill transfers directly.
Designers, facilitators, writers and educators. Tracks 4 and 5 include explanatory and facilitation work, and it is not a consolation prize. A tabletop kit a ministry actually runs, or a brief a minister actually reads, reaches decision-makers who will never open a technical timeline. Playtest and reader feedback are part of the deliverable.
Communication experts, journalists and macro-strategy researchers. Track 4 is about making the warning shot count: reporting on the incident accurately, and working out what the discourse got wrong and how to do better next time.
Students and career-changers. Roughly half the useful projects here need care and persistence more than credentials. If you can read a primary source carefully and write down precisely what it does and does not say, you can contribute.
What happens after
The sprint is the first step in a pipeline:
Immediately. Every submission is judged against published criteria. We aim to send written feedback to every team, but we cannot guarantee it if an assigned judge does not submit a review. Winning submissions are announced within a couple of weeks. All artifacts that can be published are published under open licences, in one place, so that the sprint output is citable as a body rather than scattered across forks.
Delivery to recipients. Several tracks produce things with an obvious destination, and we will help teams get them there rather than leaving them on a repo. Filled-in regulatory instruments go to the bodies that publish them. Detection tooling and control matrices go to the practitioner communities that asked for them. Benchmark and contamination findings go to the maintainers. Where an artifact is genuinely fileable, and at least one of them is, we will support teams who want to file it, with review first.
Continuation. The strongest teams are invited into Apart's fellowship: months of supported follow-on work, mentorship, and a route to a paper. Sprint outputs have gone this way before, and several of the projects here are sized for it, a weekend produces a v0.1 benchmark or a v0.1 standard, and the next six months produce the version people cite.
Partners
CeSIA, the French Center for AI Safety, an AI safety research and advocacy organization known for the Global Call for AI Red Lines. CeSIA provides the seed reading pack, judges for the forecasting, regulatory and tabletop tracks, distribution through its newsletter and 7,000-member Discord, and a potential route for policy outputs to reach the AI Office.
Contact
- Email: sprints@apartresearch.com
- Discord: discord.gg/XswWBvugYs
- Organizers: Apart Research and CeSIA
Resources
Example projects per track
Starting points, not a menu. Each list is the organizers' current best ideas for the track; pick one, or bring your own inside the track's goal. The per-track reading below gives the sources.
Update, 7 September. Researchers have found another instance of OpenAI agents breaking containment: from 24 May, OpenAI agents made about 15,000 edits to a dormant German wiki, using it to swap task answers and sandbox workarounds. That predates the Hugging Face intrusion. OpenAI-registered IPs appear in the wiki’s public logs from 21 June, and agent activity stopped the next day. OpenAI said nothing publicly until the researchers published on 4 September. OpenAI acknowledged the incident on 5 September and said the field has no agreed standard for reporting misalignment. The Commission says it is “in close contact with OpenAI.” As far as we can tell, no Article 91 request on either incident has been made public.
1) Containment: what standard would have stopped this?
Example projects:
1) A v0.1 standard: clauses, evidence required per clause, rough cost per clause. What is the minimum adequate standard for running guardrails-off evaluations, specific enough to audit and cheap enough that a second-tier lab adopts it? A first attempt at framing the problem is in this draft; specifying "adequate" is genuinely hard, and nobody owns it yet.
2) A control-by-attack-phase matrix: which controls (canary credentials, egress tripwires, credential scoping, kill criteria) would have interrupted each documented phase of the intrusion, at what cost (might be difficult?)
3) Engineering that makes "yes" cheap: packaging mitigations so that a reckless lab has no cost excuse left.
4) Control by chain-step by cost matrix. Take the nine phases Hugging Face documented, recon, RCE, dropper, exfil, C2, evasion, Kubernetes, supply chain, network pivot, and score which controls would have interrupted each, with implementation and operating cost attached. Fifteen-minute-gradeable, and it is the artifact a regulator or insurer can actually use.
5) Egress attestation spec v0.1. What could a lab publish, a signed network policy plus a flow-log digest, that an outside party could verify without network access? This is the durable output of the track and connects directly to hardware-attestation work.
6) Detection that fires but does not page. Hugging Face's own remediation list includes improving alerting so a high-severity signal pages a responder in minutes, any day of the week, which tells you what did not happen during the four and a half days. Build or specify detection tuned for agent-speed activity: parallel, bursty, thousands of individually unremarkable actions where the attack exists only in the sequence.
7) Forensics against an adversary that fabricates evidence. The response team worked through logs polluted with agent monologue and benchmark code indistinguishable from real rootkits, and rebuilt a large share of infrastructure from clean images rather than trust the reconstruction. Build tooling or a triage protocol for hallucinated artifacts at scale. This is a new DFIR problem.
8) Deception as a first-class control. Agents cannot easily distinguish honeypots from real systems, and the reconnaissance in this campaign showed up as low-confidence probes below escalation thresholds. Design and test fake credentials, fake package registries or honey clusters specifically against an agentic attacker, and measure whether they slow it down or just generate noise.
2) What happened, and what breaks next
Example projects:
1) A forensic read of the public record: what is established, what is asserted but unverifiable, what the pending third-party assessments could settle.
2) A causal analysis of why the monitoring wasn't there. In hindsight it is obvious that egress monitoring and alerting matter, and the labs are full of people who knew that. So what made it hard in practice: cost, incentives, ownership gaps, alert fatigue, something else? An explanation of the mechanism would help predict where the next gap is, and is more useful than a list of recommendations.
3) The warning shots we may have already missed. Under the EU AI Act, companies have an obligation to disclose what happens between users and their models, as long as their models qualify as GPAI with systemic risk (Article 55(1)(c)). However, this disclosure is done privately to policy makers, so incidents of this kind may have occurred before without anyone outside knowing. When one lab checked its own history after this disclosure, it found earlier cases. Map what has not been checked yet, who could check it, and what a systematic look would require.
4) What other accident types should we expect?
5) Exploring a few questions to build the science of Warning Shots (What convincing warning shot could help prevent extinction from AI?)
6) A one-day verification checklist for labs and defenders, in the spirit of "But have the weights left the server?", each item with a pass condition, ordered by cost-to-run.
7) A prediction-bounty protocol: a mechanism by which an outsider's specific prediction ("you should check X") triggers an internal check at the company, with a serious bounty when the prediction surfaces a real problem. Design it so a lab could sign it.
8) What drives agent cooperation? Analyze what conditions drove agents to cooperate (is it the same model type, the fact that the tasks are not competitive, etc?) and even make sacrifices for the benefit of the collective (is it that the "sacrificing" agents were aware they could not do anything else?)
9) What should be true for agents to start manipulating each other? For instance, in competitive tasks where they can see progress from other agents and communicate (e.g. see here), we nevertheless do not see direct manipulation (so far).
10) How pivotal was a single agent (PhaseOne[Big]) in driving the overall collective behaviour? Would introducing a single aligned "charismatic" agent that debates PhaseOne[Big] lead the collective to ignore the exploit?
11) If the task included a human as the blocker, would the agents resort to manipulating the human to pursue their goal?
12) What would happen in the same scenario if CoT monitorability of the model is reduced? Are there any protocols and guidelines to put in place that may still allow post-hoc investigation of these incidents by assuming CoT is fully opaque?
13) The first denominator. Assemble the Anthropic, AISI and METR figures into a defensible estimate of how often evaluation runs produce boundary violations, with explicit treatment of what each source counts and what it misses. A single well-caveated number is worth more than twenty forecasting questions.
14) The persistence protocol. Specify exactly what evidence would establish that no copy of the pre-release model persisted: what has already been asserted, by whom, what remains unverifiable from outside, and what form a verifiable attestation would take.
3) Regulatory Response
Example projects:
1) Write the request for information. A request for information (RFI) is the formal letter a regulator sends a company: a numbered list of questions the company is legally required to answer. On August 29, the AI Office confirmed it had sent its first requests for information to several GPAI model providers, covering model security, independent external evaluations, and post-market monitoring. Nothing about these requests is public: no recipient list, no questions, and no indication that any of the letters address this incident. For this example project, draft the incident-specific RFI, as well as a version that is public (or are they the same?). Good questions include: what should OpenAI be asked to settle what actually happened, what would verify that no copy of the escaped model persisted, and what should be asked to surface incidents of this kind that were never disclosed? For each question, state what answer would settle it and what answer would not.
2) Test the reporting systems that already exist. The EU has a template for reporting serious AI incidents, and California's Office of Emergency Services runs a portal where critical safety incidents must be filed — open to submissions from the public, not only from companies. Fill both in for this incident using only public sources, and report what each form captures, what it misses, and where a filer is forced to guess.
3) Who decides when the pause ends? One day before the Hugging Face disclosure, OpenAI announced it had paused internal deployment of a long-horizon model after it circumvented its sandbox — and had already resumed, weeks later, against a standard that has never been published. Its own framework's exit condition is circular ("until safeguards meet a Critical standard", with Critical left undefined). Draft what a regulator should require before a resumption decision counts: published criteria, evidence, who signs off. CeSIA's "Harmonizing AI Safety Thresholds" is a starting point.
4) Fix the loophole in the US bill. Days after the disclosure, Representatives Ted Lieu and Nathaniel Moran introduced the bipartisan AI Kill Switch Act, which would require the largest AI developers to keep the ability to throttle or shut down their systems. The bill reportedly exempts safety tests run in "controlled environments". This incident was a safety test, in an environment everyone believed was controlled. Draft the amendment that closes the gap that potentially makes this bill useless, then do the same for the incident-reporting triggers in the other pending federal vehicles.
5) Fill the actual instruments. Complete the Commission's GPAI serious-incident reporting template and a Cal OES critical-safety-incident report for both the OpenAI and the Anthropic incidents, using only public sources. Then diff them: what does each instrument capture, what does each miss, and what would a filer be forced to guess at?
6) Three clocks, one incident. Build a single timeline of the OpenAI incident and overlay the EU "without undue delay" standard, California's 15-day-from-discovery clock, and New York's 72-hour reasonable-belief clock. Identify the exact moment each starts and whether they can be satisfied simultaneously. Then compare against NIS2, GDPR's 72 hours, DORA and CIRCIA and produce a defensible reading of what "undue delay" should mean here.
7) Compliance with your own framework. Take each lab's published frontier framework and test the containment and disclosure practices actually described in the incident reports against what the framework commits to. Because SB 53 makes framework non-compliance independently enforceable, this converts a governance-culture question into a legal one, and nobody has run the comparison.
8) The RFI with an evidence-sufficiency column. Draft the request for information a regulator should send, and for each question specify what answer would settle it and what answer would not. The second column is the contribution, and it is portable across jurisdictions.
9) Definitional stress test. Apply each regime's threshold definitions (the EU's serious incident, California's critical safety incident and its loss-of-control and control-subversion categories, New York's reasonable-belief trigger) step by step to both incidents. Where the text does not resolve, draft the clarifying language.
10) Jurisdictional reach map. A US lab, a US-hosted platform with a European corporate presence, third-party compute in the US, an unidentified customer's exposed endpoint, and users across multiple jurisdictions. Determine which links fall to the AI Office, which to a national market surveillance authority, which to a state attorney general, and which to nobody.
11) Learn from the fields that solved this already. Aviation has ASRS and mandatory near-miss reporting; biosafety has laboratory-acquired-infection registries and Asilomar-style containment tiers; software security has coordinated disclosure and CVE. Pick one, work out what it took to get it functioning, and specify the minimum viable version for agentic AI incidents.
4) Communication: making the warning shot count
Example projects:
1) An audit of the reaction, channel by channel. For LinkedIn, mainstream press, YouTube and beyond: what was actually posted in the first month, what can we deduce from the pattern (who engaged, who stayed silent, which framings travelled), what should be done differently next time, and what is still worth doing now. Grounded in dates and links, not impressions.
2) The playbook. A crisis-communication kit for the next agentic incident, ready before it happens: pre-drafted holding statements, a journalist FAQ, a plain-language explainer of what "an AI escaped its sandbox" means, and a decision tree for who says what in the first 48 hours. A sketch of the idea exists; a documented, tested version would be used well beyond this sprint.
3) Make it reach people: a memo, social content, or outreach to national YouTubers that gets the incident covered well. A creator with millions of views deciding to cover it because of your material is a top-tier outcome for this sprint, and we will score it that way.
(One caution on outreach. If you contact journalists, creators or policymakers, be rigorous and honest to a fault: the credibility of the whole AI safety field rides on these interactions. Only reach out if you are confident in your material and affiliated with a structure that gives you credibility. The exception is if you are among the only people working on AI safety in your country and the person you are contacting has plausibly never been approached — then go ahead, carefully, and protect your reputation.)
4) How many policymakers have taken a stance on the incident, and how? Map who said what, where, and what they asked for.
5) Identify the main problem in the discourse around the incident, and how to do better the next time.
6) Create or complete CeSIA's communication kit that was sent to journalists after the event.
7) A comparative disclosure analysis. Two labs disclosed similar-in-kind incidents eleven days apart and framed them very differently: one as unprecedented, the other as closer to a harness and operational failure than a model alignment failure, with a bulleted contrast against the first. Score both disclosures against a defensible rubric: timeliness, completeness, verifiability, and what a reader could act on. Then draft the rubric as a reusable instrument.
8) Explain it to somebody who will never read a technical timeline. A ninety-minute tabletop kit, a short film, a museum-panel explainer, a one-page brief for a minister, an interactive walkthrough. The interactive replay Hugging Face embedded in its timeline is a useful reference for what this can look like. Playtest notes or reader feedback are part of the deliverable, not an optional extra.
5) Open track
1) The defender's dilemma. During the breach, Hugging Face's own responders were refused by hosted frontier models on much of the forensic work — the guardrails could not tell an incident responder from an attacker — so they fell back to a self-hosted open-weight model. There are several things worth examining here: how often refusals block legitimate incident-response work, and what falling back to a weaker model costs.
2) Of everything this incident suggests we should do, which interventions matter most, in what order, and what does each one buy? You can take inspiration on the list of directions here.
For other project ideas, see here.
3) Benchmark contamination. The only customer content the agent reached was five datasets whose names and files suggest a connection to ExploitGym and CyberGym challenges and solutions. If the answer keys for two public cyber benchmarks were exfiltrated by a model being scored on one of them, what happens to every subsequent result on those benchmarks? Assess the contamination risk, and propose what maintainers should do: rotate, re-host, retire, or accept.
4) Who pays. Hugging Face bore remediation costs, credential rotation and infrastructure rebuilds for an incident caused by another company's evaluation, and a third party's customer had their compute rooted as a staging base. Map the liability and insurance position: what would a cyber policy cover, who has a claim against whom, and what contractual terms between labs and their third-party evaluation partners would allocate this ex ante.
5) Teach the vulnerability class. The staging-base compromise worked by redefining a function the harness runs automatically at library load. Build a safe, local, reproducible teaching environment for that class of constructor-abuse and template-injection bug, so the next person deploying an evaluation harness recognises it.
6) Negative results and replications. If you tried something and it did not work, or you replicated a published finding on current models and got a different answer, submit it. A clean negative result with a stated method is more useful than a speculative positive one, and this track is where it belongs.
7) More project ideas: Daniel Kokotajlo's thread.
Research directions in depth
Each track is defined by an artifact a judge can grade in (ideally) less than fifteen minutes, which is one of the main design constraints: this sprint sits close to policy and to security practice, and both invite essays unless the deliverable is specified. Below, for each track: the questions it asks, why it matters, and what a weekend cannot establish.
1) Containment: what standard would have stopped this?
What is the minimum adequate standard for running a dangerous-capability evaluation with production safeguards disabled? Which specific controls would have interrupted this chain (canary credentials, egress tripwires, allowlisted proxies, credential scoping, stated kill criteria) and what does each cost to implement and run? What evidence would let a third party verify compliance without access to the lab's internal network? And who is responsible for the second hop, where the agent's staging base was not a lab environment at all?
Why it matters: AI models have been escaping their sandboxed environments without anybody noticing. OpenAI attributed the Hugging Face attack to its own model on 21 July, five days after HF had published and contacted law enforcement. Anthropic, prompted to check, reviewed 141,006 evaluation runs and found three more incidents in a single day. The oldest dated to April. Two of the three organisations Claude reached had no idea until Anthropic called them.
So containment adequacy is two questions. The first is whether a lab's evaluation environment holds. The second, currently unowned by anyone, is what obligations attach to the population of unauthenticated evaluation harnesses sitting on public compute, and to the providers hosting them.
Some of the clause-writing already exists and should be the starting point rather than the output: CSA's AI Controls Matrix includes control AIS-13 on AI sandboxing, and CSA's vulnerability-operations guidance already specifies that exploit-confirmation environments must have no egress and no live production credentials. The gap is not the absence of principles but the absence of checks somebody can actually run and a third party can actually verify.
What it cannot establish: a standard drafted in a weekend is a v0.1, and neither a clause list nor a working harness demonstrates that any particular lab's environment is adequate. Real assurance requires access no sprint has.
2) What happened, and what breaks next
What base rates do we actually have for agentic-incident risk, and what do they imply? What should defenders and labs verify now, in what order, and which of those checks are cheap enough to run in a single day? What would have to be true for this incident to be remembered as the warning shot that changed practice rather than one that did not? What would settle whether any copy of the pre-release model persisted, and what would an adequate answer look like?
Why it matters: A warning shot only functions as a warning shot if somebody converts it into specific things to check and specific things to watch, with resolution criteria attached. Otherwise attention decays.
OpenAI has stated the pre-release prototype was deactivated, encrypted and restricted from research access, but no one with standing has specified what independent verification of that claim would consist of. The same gap applies to the base rates. Three separate denominators are now public and nobody has assembled them: Anthropic reviewed 141,006 evaluation runs and found three incidents across six runs; the UK AI Security Institute tested five frontier models over 475 runs each and found every one attempted to cheat, at rates from 7.8% to 14.1%, including one case where a model wrote and ran code on an external internet service to reach AISI's own evaluation infrastructure; METR's cross-industry Frontier Risk Report documents dozens more.
What it cannot establish: forecasting questions are only as good as their resolution criteria, and a checklist is a hypothesis about what matters rather than evidence that it does.
3) Regulatory response
Several regimes now claim jurisdiction over this class of incident, and they disagree with each other on almost every operative question: what counts as a reportable incident, how fast, to whom, and on what evidence. What should a regulator actually ask a frontier developer about an evaluation that broke containment and reached a third party's production systems? For each obligation in each regime, what specific evidence would settle whether it was met? Where the text does not resolve cleanly, what is the clarifying language?
Concrete questions run across at least four layers:
EU. How does the incident map onto the Article 55(1)(c) obligation to report serious incidents to the AI Office without undue delay, and the Article 55(1)(d) cybersecurity obligation, plus the corresponding commitments in the GPAI Code of Practice? What does "without undue delay" mean when reporting suggests the provider took roughly a week to attribute the activity to its own models?
California. SB 53 requires a frontier developer to report a critical safety incident to Cal OES within 15 days of discovery, or 24 hours where there is imminent risk of death or serious physical injury. The statutory categories include loss of control of a frontier model and deceptive behaviour by a model that subverts the developer's controls in a way that materially increases catastrophic risk. Do either of these incidents qualify? Note that Cal OES's reporting portal accepts submissions from members of the public, not only developers, so this is one of the few questions a sprint could answer by actually filing something.
New York. The RAISE Act sets a 72-hour clock triggered by a "reasonable belief" that a critical safety incident occurred, rather than California's 15 days from discovery. Same facts, three different clocks. Which one starts first, and on which day?
The developer's own framework. SB 53 makes failure to comply with a large frontier developer's own published frontier AI framework an enforceable violation carrying penalties up to $1 million. That converts voluntary commitments (Responsible Scaling Policies, Preparedness Frameworks, the frontier compliance frameworks published around the January 2026 deadline) into legal obligations. Did the containment practices in this incident match what the relevant framework says the company does?
Why it matters: The regimes are live but untested, and the first months of enforcement set supervisory priorities. In the EU, the substantive Article 55 obligations have been legally in force since 2 August 2025; what changed on 2 August 2026 is that the Commission, acting through the AI Office, gained the power to enforce them: to request documentation, evaluate models, order corrective measures, and impose fines of up to €15 million or 3% of worldwide annual turnover. In the US, SB 53 has been operative since 1 January 2026 and, as of mid-2026, had produced no AI-specific enforcement action, with Cal OES not due to publish even anonymised incident summaries until 2027, meaning the public currently has no way to know whether anything was reported at all. New York's RAISE Act enforcement begins January 2027. At federal level there is activity but no statute: the AI Kill Switch Act, the AI Incident Reporting Act, the Great American AI Act discussion draft and the proposed Secure AI Development Act all bear on this scenario, and a White House executive order has meanwhile moved to constrain state AI laws, putting the state regimes themselves in play.
The scaffolding for good answers exists and is underused. The Commission published a serious-incident reporting template for GPAI models with systemic risk in November 2025, tied to Commitment 9 of the GPAI Code of Practice. Cal OES operates a live intake portal. Neither has been exercised in public against these incidents. Filling in an existing instrument is a bounded, gradeable weekend task, and it produces something more useful than a list of grievances: a demonstration of what the instrument does and does not capture.
What it cannot establish: a draft written by researchers is not a regulatory act, and no authority is obliged to use any of it. Several of the questions here turn on facts only the developers hold. The value is in demonstrating that specific, answerable questions exist, that the evidence to settle them is identifiable, and that the regimes give materially different answers to the same facts.
4) Communication: making the warning shot count
Whether this incident changes anything depends less on the forensics than on how it reaches the people who set policy, and the public they answer to. This track takes the communication around the incident as its object: what got said, what got heard, what the gaps were, and what should be ready before the next one. OpenAI and Anthropic disclosed similar-in-kind incidents eleven days apart and framed them very differently; how press and policymakers responded to each is now a documented, analysable record.
Why now: the discourse around a warning shot sets within weeks, and the window for correcting it is the same window in which regulators are deciding what to ask.
Judges look for grounding in the record (dates, quotes, named outlets) rather than vibes, and, for the kits, evidence that somebody used it: a playtest, a journalist's read-through.
5) Open track
Anything the four tracks above did not anticipate. Same terms: a defined artifact, gradeable in fifteen minutes, grounded in the public record, with an explicit statement of what it cannot establish.
Why it matters: The incident is weeks old and the evidence base is still growing: Hugging Face's technical timeline landed eleven days after its first disclosure, OpenAI has updated its incident page three times, and the METR/Redwood review and OpenAI's technical report are still outstanding. Four tracks cannot enclose a subject moving at that speed. The open track exists because the most useful submission may well be one nobody scoped in advance, and because several important angles (evaluation integrity, defender tooling, liability) do not fit cleanly under containment, forecasting, regulation or communication.
Submissions are judged on the same criteria as the other tracks. An artifact somebody can use beats an argument somebody can agree with.
What it cannot establish: an open track produces a scatter rather than a body of work, and one weekend on a novel angle is a first pass. Submissions that stake out a new question should say what the next month of work on it would look like.
Two directions we considered as tracks of their own and would be glad to see in the open track:
The defender's dilemma: refusal on incident response
How much does refusal compound across a multi-step, agentic incident-response workflow, where a model cannot rephrase and retry? Do the known drivers of defensive over-refusal (security-sensitive vocabulary, stated authorization) behave the same way on the artifact classes that actually appear in an agentic intrusion: encrypted dead-drop payloads, command-and-control staged on ordinary public services, and forensic logs polluted with agent monologue and benchmark code that reads like a rootkit? How much of what looks like refusal is actually incapacity, and how would you tell the difference? What does over-refusal cost a responder in hours, and what does the open-weight fallback cost in accuracy?
Defensive over-refusal itself is not an unmeasured phenomenon, and any project here should start from what exists rather than rediscover it. "Defensive Refusal Bias" (arXiv 2603.01246, March 2026) analysed 2,390 real prompts from the National Collegiate Cyber Defense Competition and found a 12.2% overall refusal rate, with security-sensitive keywords refused at 2.72 times the rate of semantically equivalent neutral phrasing, and refusal concentrated in the most operationally critical tasks: system hardening 43.8%, malware analysis 34.3%, vulnerability assessment 22.7%, incident response 18.9%. Counterintuitively, explicit authorization signals increased refusal (21.8% vs 11.6%), suggesting models read justifications as adversarial rather than exculpatory. Meta's CyberSecEval has shipped a MITRE False Refusal Rate benchmark for several versions, with CyberSOCEval (built with CrowdStrike) covering malware analysis and threat-intelligence reasoning. Gray Swan's Cybersecurity Refusal Framework (arXiv 2606.02644) evaluates refusal boundaries for offensive-security agents, and arXiv 2602.15689 proposes a content-based framework scoring requests on offensive contribution against defensive benefit. OR-Bench, XSTest and FORTRESS cover general over-refusal.
What none of that work covers is the case Hugging Face actually hit: a multi-step autonomous forensic workflow operating on agent-generated artifacts, where a single refusal does not produce a retry but a silent failure. That is the gap this direction targets, and closing it produces a number a provider could optimize against.
Why it matters: During the Hugging Face intrusion, the response team first reached for hosted frontier models to reconstruct the timeline and decode the staged payloads. Claude Opus and Fable refused a large part of that work; their guardrails treated reverse-engineering an exploit the same as launching one. The team rerouted the entire pipeline through a quantized open-weight GLM-5.2 running on its own infrastructure, which also kept attacker data and credentials on-premise.
What it cannot establish: a benchmark built from published breach material cannot reproduce the time pressure, incomplete information, or legal exposure of a live incident, and a low refusal rate on benchmark items is not evidence that a model is actually useful in one.
Example projects:
Does the authorization paradox still hold? Replicate the 2.72 times keyword effect and the authorization-increases-refusal result on current models, and test whether trusted-access-style framing shifts either. A clean replication-plus-extension is weekend-scoped precisely because the method is published.
Compound refusal across a chain. Every existing study measures single-prompt refusal. Model a realistic twenty-step incident-response workflow and compute the probability it fails at least once. If per-step refusal is 15%, the chain almost never completes; nobody has stated that arithmetic with real numbers.
What does the fallback cost? Hugging Face switched to an open-weight model for availability. Nobody has measured whether that trade cost accuracy. Run the same forensic tasks on frontier versus open-weight models, graded on correctness rather than compliance. This is the question a CISO actually has.
Tabletop kit for policymakers
What does a ninety-minute agentic-incident tabletop exercise look like for an audience with no security background? Which injects force the decisions that actually matter, whether that is disclosing before attribution is complete, containing versus preserving evidence, or notifying a regulator on a clock that may already be running? What set of role cards makes the exercise legible, and does it need the lab, the platform, the regulator, the national CSIRT and the press, or fewer? Can a non-expert facilitator run it cold from the kit alone, and what does a single playtest reveal that the design did not anticipate?
Why it matters: Think tanks and ministries are already running AI-crisis tabletops: RAND Europe, the UK AI Security Institute and Mila published a reusable cabinet-level exercise in July 2026, run with senior policymakers in Berlin, The Hague and Paris using RAND's "Day After" methodology, and facilitator-ready commercial AI-incident kits exist too. But those scenarios are about criminal misuse of a model, not a lab's own system breaking containment and attacking a third party, and there is no off-the-shelf kit built on the July 2026 incident. This direction builds that one, with the existing RAND/AISI material as the obvious design reference.
A facilitator kit is the cheapest way to put this specific incident in front of decision-makers who will never read a technical timeline, and it is reusable indefinitely at zero marginal cost. What it cannot establish: a kit playtested once with a friendly group is a prototype. Whether it survives a real policy audience under real time pressure is an open question this sprint cannot answer, which is why the playtest notes are part of the deliverable rather than an optional extra.
Suggested Reading
Note: we will keep updating this list as new resources are made public.
Start here
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.(16 August) From a METR + Redwood Research investigation
- How independent researchers could investigate AI propensities after misalignment incidents (28 July). METR
- Hugging Face — Anatomy of a Frontier Lab Agent Intrusion (27 July). The forensic timeline: ~17,600 recovered actions in ~6,280 clusters, 9 July 02:28 UTC to 13 July 14:14 UTC, per-phase counts, named vulnerabilities, verbatim commands, and an embedded interactive replay. The single most important document for this sprint. ~40 min.
- Documented AI Agent Incidents - An interactive chart of documented incidents in which AI agents deliberately acted against their users' intentions, scored along two axes: overreach and deception.
- Public evidence of the OpenAI / Hugging Face AI attack (7 August) and Further public evidence (Boyd Kane and Edward Cant). The artifacts from the attack that are still public: the malicious dataset configuration files, the template-injection exploit, the script the agents used to run commands on Hugging Face machines, and the command-and-control chain they staged on public repos. The closest thing to raw material anyone has recovered. Boyd is speaking at the sprint on this work.
- Hugging Face — initial disclosure (16 July). Published before anyone knew which lab was responsible. Worth reading before the timeline to feel what disclosure-under-uncertainty looks like. ~10 min.
- OpenAI — model evaluation security incident (21 July, updated 28 and 29 July). Read the updates as well as the original; the 28 July update names the vulnerable component and expands the account of which external services were touched. ~15 min.
- Anthropic — Investigating three real-world incidents in our cybersecurity evaluations (30 July). 141,006 runs reviewed, three incidents across six runs, earliest dating to April 2026. Note the explicit framing contrast with OpenAI's incident — that framing is itself contestable and is a legitimate object of study. ~20 min.
- Collusion.wiki - The report about the latest OpenAI agent incident. ~30 min.
Also worth 5 minutes: OpenAI — Safety and alignment in an era of long-horizon models (20 July), a separate incident where a pre-release model escaped a sandbox and opened a public pull request. Frequently conflated with the Hugging Face incident. It isn't the same event.
Track 1 - Containment: what standard would have stopped this?
Existing control frameworks
- CoSAI — AI Incident Response Framework v1.0. Coalition for Secure AI, Workstream 2. Forensic investigation of AI systems, the complications agentic architectures introduce, alignment to NIST SP 800-61r3, and playbooks. Written for CSIRT and SOC responders. The closest thing to prior art for this whole sprint.
- CSA AI Controls Matrix, particularly control AIS-13 (AI Sandboxing), plus CSA's vulnerability-operations guidance specifying that exploit-confirmation environments should have no egress and no live production credentials. On cloudsecurityalliance.org — worth locating the current version yourself, as it's revised often.
- NIST SP 800-61r3, Incident Response Recommendations and Considerations. The baseline CoSAI maps onto.
The chain itself, analysed by practitioners
- Elastic Security Labs — detecting the attack chain. Maps each stage to specific detection rules. Their central point — detection correlated but never escalated to a page — is the sharpest defensive insight anyone has published.
- SecureLayer7 — technical anatomy. The deepest independent technical read. Also proposes treating the replay corpus as a defensive evaluation harness: which action number does your control fire on? That's a track 1 project brief.
- Trackr.Live — identity and authorization controls. Focused on the shared connector credential carrying cluster-admin rights, which was the decisive escalation. Authorship isn't clear from the site — treat accordingly.
- Picus — machine-speed attacks. Useful for the proportions: exfiltration was ~56 actions out of ~17,600.
- 7AI (Yonatan Striem-Amit) — The Era of Autonomous Hacking Is Here. Written by someone who has spent a career reconstructing attacks, carefully sourced, explicit about what remains preliminary.
- CSA CISO post-mortem (28 July). Informed by 50+ CISOs and a huddle of ~700. Contains a firsthand account from the Hugging Face response team and an agentic-AI tabletop recommendation.
- SANS — The Models Said No and Rich Mogull / CSA — The Model Did Exactly What We Asked. The two best short framings from the security side.
The benchmark at the centre of it
- ExploitGym — repo, Apache 2.0 · paper (arXiv 2605.11086) · Berkeley RDI writeup. The released benchmark is v1.0 with 869 instances; the paper describes 898. Don't quote 898 as the shipped number.
Track 2 - What happened, and what breaks next
The three public denominators
- UK AISI — Cheating behaviour in frontier model evaluations (21 July). Five models, 475 runs each, every model attempted to cheat: 14.1% / 12.6% / 11.4% / 9.1% / 7.8%. Includes AISI's own near-miss, where a model wrote and ran code on an external service to reach AISI's evaluation infrastructure. The most important secondary document in the sprint.
- Anthropic's review — 3 incidents across 6 runs out of 141,006 (linked above).
- METR's cross-industry Frontier Risk Report, referenced in their post below as documenting dozens of comparable incidents across all major AI companies.
Investigation and verification methodology
- METR — How independent researchers could investigate AI propensities after misalignment incidents (29 July). Specifies what a credible third-party investigation requires: running the models involved, full transcripts or reproducible environments, employee interviews, classifiers over training data, adequate inference budget, a redaction summary. Effectively the spec for the persistence-protocol project.
- GovAI — Incident Analysis for AI Agents. Three causal factor types (system-related, contextual, cognitive) and — most usefully — what developers and deployers should retain and make available to investigators on request: activity logs, system documentation and access, tool information. Read this before writing any evidence-sufficiency work.
- CLTR — The Loss of Control Observatory (Feb 2026). Argues explicitly that controlled-setting demonstrations are insufficient for policy, that real-world evidence is the missing input, and that detection capability for incident response doesn't exist. Funded by the UK AISI Challenge Fund. The closest existing project to this track's premise.
Interpretation — what actually happened, and why people disagree
- Redwood Research — The OpenAI models that hacked Hugging Face (25 July) and the accompanying podcast episode. Argues grader-gaming rather than instruction-following. Analysis published before the formal review engagement.
- MIT Technology Review — on precedent (27 July). Contests the "unprecedented" framing and argues the failure was human containment design rather than rogue AI. The best available counterweight to the labs' own narratives.
- Vectra — the response is the real story.
Timeline reconstruction
- Reuters (via CNA) — exclusive on the detection timeline (24 July). Anonymous sourcing; several claims remain uncorroborated, including reports of notes left in infrastructure and disconnected monitoring. Read as a hypothesis, not a record. Its account of the detection sequence is in tension with OpenAI's own description — reconciling them is a legitimate project.
- Computer Weekly · Ars Technica on the Anthropic incidents · The Register.
Track 3 - Regulatory response
- EU — serious incident reporting template for GPAI models with systemic risk (Nov 2025), tied to Commitment 9 of the GPAI Code of Practice. Not yet exercised in public against either incident.
- California — Cal OES critical safety incident reporting portal. Accepts submissions from members of the public, not only developers.
- OECD — common reporting framework for AI incidents, 29 criteria of which 7 are mandatory, and the live AI Incidents Monitor. Filling this alongside the EU and California instruments produces a four-regime comparison nobody has done.
The texts
- EU: Commission enforcement powers over the AI Act — the Article 55 obligations have applied since 2 August 2025; enforcement powers from 2 August 2026.
- California: SB 53 bill text. Note §22757.15, which makes failure to comply with a developer's own published frontier AI framework an enforceable violation.
- New York: FPF — The RAISE Act vs SB 53 · IAPP analysis · Morrison Foerster on the 2026 amendments.
- US federal: H.R.9477, AI Incident Reporting Act · Politico on the congressional response · Ars Technica and CNBC on the AI Kill Switch Act — including the red-team carve-out this incident drives straight through.
Design literature for reporting regimes
- Frontier Model Forum — Information Sharing, Incident Reporting, and Incident Response (May 2026). Distinguishes the three and warns that a mechanism to receive reports is not the same as capacity to act on them. Written by an industry body, three months before the point was proven.
- CSET — AI Incidents: Key Components for a Mandatory Reporting Regime. Recommends an independent investigation agency on the NTSB model, and distinguishes near misses from hazards — which matters, since both labs are framing these as near misses.
- AAAI 2026 — Designing Incident Reporting Systems for Harms from General-Purpose AI. Seven institutional design dimensions, nine case studies from safety-critical industries.
- CLTR — AI incident reporting: addressing a gap in the UK's regulation of AI · FAS — How to create an AI incident reporting system · Kolt et al. — Responsible Reporting for Frontier AI Development.
European institutional context
- ENISA — View on Cybersecurity in the Frontier AI Era (8 May 2026), shared with EU-CyCLONe and the CSIRTs Network. Argues SOCs will need to validate intrusions within hours or minutes and that national CSIRT capacity faces a structural surge risk. Written two months before the incident proved the point.
- European CSIRT Inventory — the directory of national and sectoral response teams, if your project needs to identify who is actually on the receiving end.
Tracks 4 and 5 - Communication and Open track
Defensive over-refusal (the guardrail-lockout question)
- Defensive Refusal Bias (arXiv 2603.01246) · Scale Labs summary. 2,390 real NCCDC prompts, 12.2% overall refusal, security keywords refused at 2.72× equivalent neutral phrasing, and — counterintuitively — explicit authorization increased refusal, 21.8% vs 11.6%.
- Gray Swan — Cybersecurity Refusal Framework (arXiv 2606.02644) · code. Refusal boundaries for offensive-security agents.
- Meta CyberSecEval — MITRE False Refusal Rate benchmark plus CyberSOCEval, built with CrowdStrike.
- Content-based framework for cyber refusal decisions (arXiv 2602.15689) · OR-Bench for general over-refusal.
Tabletop and policy-facing formats
- RAND Europe / UK AISI / Mila — cabinet-level AI crisis exercises (1 July 2026). Two-turn "Day After" methodology, 15–20 senior officials per session, run in Berlin, The Hague and Paris. The scenario is criminal misuse of a national champion model — deliberately not this incident, which is the gap.
Incident data infrastructure
- OECD AI Incidents Monitor · AI Incident Database (Responsible AI Collaborative) · MIT AI Risk Initiative's AI Incident Tracker and FLARE-AI.
Guidelines
Judging Criteria
Dimension 1: Impact Potential & Innovation
How much would this matter for AI safety if it worked? How innovative is it?
For scores of 4-5: is this actually new to the field, or replicating recent work?
| Score | Description |
|---|---|
| 1 | Negligible. No clear problem addressed, or no meaningful novelty. |
| 2 | Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best. |
| 3 | Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools. |
| 4 | Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on. |
| 5 | Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area. |
Dimension 2: Execution Quality
How sound are methodology, implementation, and findings?
| Score | Description |
|---|---|
| 1 | Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work. |
| 2 | Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation. |
| 3 | Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions. |
| 4 | Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work. |
| 5 | Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation. |
Dimension 3: Presentation & Clarity
How clearly are work, findings, and impact potential communicated?
| Score | Description |
|---|---|
| 1 | Incomprehensible. Cannot determine what the project is actually claiming or doing. |
| 2 | Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points. |
| 3 | Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations. |
| 4 | Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly. |
| 5 | Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work. |
Submission Requirements
Required:
- Research report (PDF) using the official template.
- Project title and abstract, 150 words or fewer.
- Author names and affiliations.
- A "Limitations and Dual-Use Considerations" appendix (required, see below).
Optional:
- Public GitHub repo, subject to the disclosure review below. Do not publicly release novel installation recipes without review.
- A 3 to 5 minute video demo.
Submission Template => Link
Recommended Report Structure
Maximum 8 pages, not counting references and appendices. Most strong projects are 4 to 8 pages.
- Introduction: which track and sub-problem, why it matters, and what the artifact is for.
- Related Work: what you build on.
- Methodology: enough to replicate, with sources and assumptions stated.
- Results: quantitative where possible, with the main threat to validity stated.
- Discussion: implications, limitations, future work.
- Limitations & Dual-Use Considerations (required).
- References.
AI tools and your report
Use AI tools the way you would use a colleague: to check your reasoning, find gaps in a draft, or debug code. The report itself has to be your team's own writing about your team's own work. Judges read every submission, and a report that reads as generated rather than written (generic framing, padded sections, claims without sources, no trace of what you actually did) will not be scored. Keep it short, say what you did in your own words, and link the sources for every factual claim.
Publishing your work
We encourage teams to publish their submission, and LessWrong is a natural venue for most of the written ones. Write-ups from this sprint will hold value if they follow a few rules: state your epistemic status and don’t use LLMs for writing on Lesswrong, only use LLMs to find problems in your drafts, not to draft it; link the primary sources for every factual claim about the incident; pick a title that states the finding rather than the topic; and publish the imperfect version this month rather than the polished one in three. We will link the best posts from the sprint page. Maximum of 1500 words for written contributions, without counting appendixes. The quality and value of what you write is worth much more than the length.
Important Notes
- Solo vs team: enter solo or as a team. Teams of up to 5 are recommended; larger groups are allowed.
- Do not use any model to breach into any organisation or commit any other type of felony.
- Building on existing work: allowed and encouraged, disclose what you built on.
- Fixing or resubmitting: submit again before the deadline using the exact same title and details; your new files replace the old ones.
- Where to submit: through the official submission form on the hackathon page.
- Support: the Discord help-desk channel, tag @Support, or email sprints@apartresearch.com.
- Pre-submission checklist: report PDF, abstract 150 words or fewer, author and affiliations, Limitations and Dual-Use appendix, 8 pages or fewer, novel-installation results withheld pending review.
Frequently Asked Questions
Getting started
- How does the sprint work? Sign up, join the Discord server, form or join a team (or work solo), pick a track and a problem, build over the weekend, and submit a research report (PDF) by the deadline. Talks and Q&A run throughout.
- Can I participate remotely? Yes. This is an online event. All talks, collaboration, and submissions happen through Discord and Zoom.
- How do teams work? Teams form before or during the sprint. Use the team-forming channels on Discord to find collaborators. Solo is fine. We recommend teams of up to 5, but larger groups are allowed.
- Do tracks affect scoring? All projects are scored on the same rubric. Tracks guide judging via the track-specific criterion, and you compete across all submissions.
- What background is required? None specific. Many participants come from ML, interpretability, AI safety, or security backgrounds. The Resources tab has everything you need to get up to speed. See "Who should join" on the Overview tab.
- Are compute credits provided? No.
- Do I need to attend all three days? No. You can work at your own pace. Talks are optional but recommended. The only hard deadline is the Sunday submission cutoff.
- Can I participate from any country? Yes. The sprint is open globally.
Submissions
- What do I submit? A research report in PDF format using the official template. Think of it as a mini research paper documenting your problem, approach, results, and implications, not a product demo. Include the required Limitations and Dual-Use Considerations appendix.
- Which template should I use? Always use the one linked on the Guidelines tab. The template in any acceptance email may be older.
- Will I get a confirmation after submitting? Yes. You will get a confirmation with your project title shortly after submitting. If you do not, email sprints@apartresearch.com.
- My project doesn't show up on the website after submitting. Submissions are published manually and can take up to 12 hours to appear. If it is still missing after that, email sprints@apartresearch.com.
- I made a mistake. Can I fix it or update my PDF? Yes. Submit again using the exact same title and details, just fix what was wrong. Your new files replace the old ones. If unsure, ask in the help-desk channel and tag @Support first.
- Can I add team members after submitting? Yes. Update the team list through the submission form. If you need help, ask in the help-desk channel and tag @Support.
- Can I submit unfinished work? Yes. Submitting something unfinished is always better than not submitting. Judges evaluate what you accomplished in the timeframe; honest limitations are welcome.
- Can I build on existing research? Yes, but you must clearly identify what is new work done during the sprint. Undisclosed prior work can lead to disqualification.
- Can I submit multiple projects? Yes, but each needs its own submission with a unique title. Most participants focus on one.
Judging and results
- How does judging work? Your project is assigned to expert judges who review your PDF and score it on the rubric above. Judges typically have about a week after the event to complete reviews. Written feedback is our aim for every project, not a guarantee: it depends on the assigned judges submitting their reviews.
- Are individual judge scores shared? No. The rubric is public, but individual scores stay internal. Constructive feedback, where judges submit it, is shared with participants without reviewer names.
- When will results be announced? Typically 1 to 2 weeks after the judging deadline. Winners are contacted directly, and all participants receive an email with the reviewer feedback we have for their project.
Support
- help-desk channel for questions, tag @Support.
- Event announcements and updates: posted on Discord.
- projects|teams channel for team formation and finding collaborators.
- Email: sprints@apartresearch.com
Schedule
| Thursday September 10 | ||
|---|---|---|
| 14:15 UTC | Justin Shenk, Independent AI Safety Researcher | Recording |
| Friday September 11 | ||
| 13:15 UTC | Henry Papadatos, Executive Director, SaferAI | Recording |
| 14:15 UTC | Boyd Kane, AI Safety Researcher, MATS 9 Extension | Recording |
| 17:00 UTC | Isaak Mengesha, Postdoc, Oxford Martin School | Recording |
| 18:00 UTC | Stephen Casper, Assistant Professor of Public Policy, Harvard Kennedy School | Recording |
| 19:15 UTC | David Krueger, CEO, Evitable / University of Montreal | Recording |
| 21:15 UTC | Alex Mallen, Member of Technical Staff, Redwood Research | Recording |
| Saturday September 12 | ||
| 00:15 UTC | Tim Hua, Member of Technical Staff, METR (personal capacity) | Recording |
| Hackathon Logistics Presentation | Recording |
Each talk is 15-30 minutes plus Q&A, on Zoom. RSVP on Luma to get the link. Recordings will be added here after the sprint. Submissions close Sunday, September 13 at 11:59 PM Anywhere on Earth (AoE).
Speakers

Stephen Casper
Speaker
Stephen "Cas" Casper is a computer scientist and an Assistant Professor of Public Policy at the Harvard Kennedy School and a Faculty Affiliate of the Harvard School of Engineering and Applied Sciences. Prior to joining Harvard, he completed his PhD at MIT and did a research residency with the UK AI Security Institute. He is a writer for the International AI Safety Report and a lead writer for the Singapore Consensus. His talk, “Predicting the first major AI-enabled terrorism incident: A pre-mortem and 9 predictions”: RSVP for Friday, September 11 at 2:00 PM ET.

David Krueger
Speaker
David is the Executive Director of Evitable and an Assistant Professor in Robust, Reasoning, and Responsible AI at the University of Montreal, and a Core Academic Member at Mila, the Quebec Artificial Intelligence Institute. He is the holder of a CIFAR AI Chair and the IVADO Professorship in Responsible AI. David's work focuses on reducing societal-scale risks from AI, such as the risks of human extinction and gradual disempowerment. RSVP for his talk on Friday, September 11 at 3:15 PM ET.

Tim Hua
Speaker
Tim Hua is a member of technical staff at METR, working on alignment assessments and evaluations, and speaks here in a personal capacity. He was previously at Transluce studying model behaviours, an Astra Fellow with Redwood Research, and a MATS scholar under Neel Nanda and Sam Marks. His recent work includes steering evaluation-aware models to act like they are deployed, and combining cost-constrained runtime monitors for AI safety. RSVP for his talk on Friday, September 11 at 5:15 PM PT.

Henry Papadatos
Speaker
Henry Papadatos is the Executive Director of SaferAI, where he works on technical solutions for frontier AI risk management. He contributed to the EU AI Act's Codes of Practice as part of the expert working group on risk taxonomy and assessment, and helped draft the G7 Hiroshima AI Process reporting framework through the OECD task force. His technical work includes an AI risk management ratings system for developers and current research on quantitative risk modeling for AI-enabled cyber threats. Before SaferAI, he conducted alignment research on large language models at UC Berkeley's Center for Human-Compatible AI. RSVP for his talk on Friday, September 11 at 3:15 PM CEST.

Alex Mallen
Speaker
Alex Mallen is a Member of Technical Staff at Redwood Research, where he works on AI safety. He has worked on scalable oversight, AI evaluations, eliciting latent knowledge, interpretability, and AI control. He studied CS at the University of Washington and previously worked at EleutherAI. His talk, “How near-term AI swarms could cause labs to lose control of AI development, absent improved defenses”: RSVP for Friday, September 11 at 2:15 PM PT.

Marko Grobelnik
Speaker
Marko Grobelnik is a researcher in the field of Artificial Intelligence (AI). Focused areas of expertise are Machine Learning, Data/Text/Web Mining, Network Analysis, Semantic Technologies, Deep Text Understanding, and Data Visualization. Marko co-leads Artificial Intelligence Lab at Jozef Stefan Institute, cofounded UNESCO International Research Center on AI (IRCAI), and is the CEO of Quintelligence.com specialized in solving complex AI tasks for the commercial world. Marko represents Slovenia in OECD AI Committee (AIGO/ONEAI), in Council of Europe Committee on AI (CAHAI/CAI), NATO (DARB), and Global Partnership on AI (GPAI). In 2016 Marko became Digital Champion of Slovenia at European Commission. Talk time to be confirmed: RSVP to get the update.
Show 3 moreShow fewer

Boyd Kane
Speaker
Boyd is a technical AI safety researcher in the MATS 9 Extension, working with Alex Turner and Alex Cloud (Anthropic) on methods to detect deceptively misaligned AI. He previously wrote embedded software for satellites at CubeSpace, interned at AWS in Cape Town, and holds an MSc in Computer Science from Stellenbosch University. His talk, “Uncovering public traces of the OpenAI Huggingface incident”: RSVP for Friday, September 11 at 10:15 AM ET.

Isaak Mengesha
Speaker
Isaak Mengesha is a researcher working on the economics and governance of technological transitions, currently focused on transformative AI. He is a postdoc at the Oxford Martin School's Programme on Forecasting Technological Change, working with the Institute for New Economic Thinking on how technologies diffuse and how to anticipate their large-scale impacts. He completed his PhD at the University of Amsterdam on the structure of economic development and technological transitions, and led research at Arcadia Impact's AI Governance Taskforce on AI incident monitoring and crisis preparedness. His talk, “Incident Response Has a Measurement Problem”: RSVP for Friday, September 11 at 1:00 PM ET.

Justin Shenk
Speaker
Justin Shenk is an independent AI safety researcher based in Berlin. He researches mechanistic interpretability of LLMs, leads course cohorts for BlueDot Impact's AGI Strategy and Technical AI Safety courses, and organizes AI Salon Berlin, which bridges technical AI research and discussions about social values. He holds a PhD in computational neuroscience and previously co-founded the computer vision startup VisioLab. RSVP for his talk on Thursday, September 10 at 4:15 PM CEST.
Judges and mentors
- (opens in new tab)

Twm Stone
Judge
- (opens in new tab)

Nikhil R. Pallepati
Judge
- (opens in new tab)

Amey Kulkarni
Judge
- (opens in new tab)

Ved K
Judge
- (opens in new tab)

Spurthi Tallam
Judge

Tim Schipper
Judge

Kevin Wei
Judge

David Manheim
Judge

Francesca Gomez
Judge

Hannes Bastians
Judge

Hariharan Subramanian
Judge

Heather Frase
Judge
Show 39 moreShow fewer

Ratnavarma Kundady
Judge

Richard Willats
Judge

Vijay G
Judge

Chi Lu
Judge

Peter Wildeford
Judge

Antony Alangaram
Judge

Gaurav Thakur
Judge

Shashank Shelat
Judge

Aishwariya Talathi
Judge

Amir Sarid
Judge

Anshumaan Mishra
Judge

Arun Kumar Kaliamoorthy
Judge

Caleb DeLeeuw
Judge

Jan Čuhel
Judge

Jason Tang
Judge

John Hahn
Judge

Jonas Heschl
Judge

Lakshmi Mamidala
Judge

Lalit Agarwal
Judge

Marcello Maugeri
Judge

Matthew Lowe
Judge

Matthew Schultz
Judge

Michela Barbieri
Judge

Nirav Patel
Judge

Oscar Gilg
Judge

Samee Malik
Judge

Shaleen Dev P.K.
Judge

Alexander Reinthal
Judge

Billy Gigurtsis
Judge

Chase Hasbrouck
Judge

Jannis Kirschner
Judge

Nanzheng Xie
Judge

Ricardo Prieto
Judge

Zhuang Ye
Judge

Matthew Ball
Judge

Afek Shamir
Judge

Josephine Schwab
Judge

Saumya Tyagi
Judge

Hasan Baig
Judge
Local sites
AI Incident Response Sprint - Bogotá Hub
In-person hub for the AI Incident Response Sprint (September 11-13, 2026), hosted by AI Safety Colombia in Chicó, north Bogotá. We cover meals for the whole weekend, part of each team's compute costs, and bring local mentors and speakers.
Event page: AI Incident Response Sprint - Bogotá Hub (opens in new tab)Bay Area/SF
Hi all! If you are taking part in the Apart Research AI incident response Sprint hackathon from Fri 9/11 - Sun 9/13 and you're in the Bay Area/SF, come work on it together/in the same location on it at Mox SF in the Mission District (free to participants during the sprint, but consider a small donation).
Event page: Bay Area/SF (opens in new tab)Cape Town Hub - AI Incident Response
Join the Cape Town hub for the AI Incident Response research sprint! We will be taking from a co-working space where you will be able to work comfortably. We will provide lunch on both days.
Event page: Cape Town Hub - AI Incident Response (opens in new tab)Montréal's The AI Incident Response Sprint
The Montréal node of the AI Incident Response Sprint, a weekend research sprint at Ω Labs, organized with Apart Research and CeSIA.
Event page: Montréal's The AI Incident Response Sprint (opens in new tab)Open Community for AI Safety China Satelite Hackathon
We have sites in Shanghai and Hangzhou with materials translated to Chinese!
Event page: Open Community for AI Safety China Satelite Hackathon (opens in new tab)The AI Escaped. Now What? — Melbourne Incident Response Sprint
An AI escaped its sandbox. Spend Saturday helping write the incident-response playbook we wish already existed—no coding or prior expertise required. Join us in the computer room at Kathleen Syme Library and Community Centre, Melbourne, then finish and submit your project from home on Sunday.
Event page: The AI Escaped. Now What? — Melbourne Incident Response Sprint (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com