-

Online & In-Person

AI Incident Response Sprint

A three-day sprint for researchers, engineers, and security practitioners to turn public incident evidence into practical response methods for autonomous AI Incidents.

17

Days To Go

Overview

Resources

Guidelines

Schedule

Overview

Arrow

In this 3-day research sprint, you will turn the first documented cases of an AI system autonomously breaching a third party into artifacts that defenders and regulators can actually use, working in teams to produce containment standards, escape-detection harnesses, forecasting question sets, draft regulatory information requests, playtested tabletop exercises or anything that will help us be more ready for the next one.
Co-organized by Apart Research and CeSIA, this sprint sits at the intersection of AI safety, security incident response, technology regulation, and forecasting. No prior background in AI incident response is required.

Cash Prizes

$2,000 in cash prizes across all tracks

🥇 1st Place

$1,000

🥈 2nd Place

$500

🥉 3rd Place

$300

🏅 4th Place

$100

🏅 5th Place

$100

Fast-track and continuation

  • Follow-up program: top teams continue through the Apart Fellowship for further research and mentorship; the timeline is shared with invitations.

  • What winners receive beyond cash: fellowship fast-track, mentor introductions from the organizing team and judges, and a pathway toward a publishable write-up.

What this Sprint is about

AI incident response is the practice of turning incidents in which an AI system is itself the actor into fewer incidents later. That spans the operational work: detecting, containing, and reconstructing what an autonomous agent did across systems it was never authorized to touch, and the strategic question a warning shot poses: which disclosure and regulatory responses actually reduce risk, and which merely suppress the visible evidence that would have prompted action while the stakes were still small.

What participants will do

Over three days you will take one unusually well-documented real-world incident and turn it into an artifact somebody outside this sprint can use.

The format. Teams of one to five people pick a track, come up with a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no specific background needed.

Research Tracks

1) Containment: what standard would have stopped this?

The models escaped through the one network egress their sandbox permitted. 

Example projects:

1) A v0.1 standard: clauses, evidence required per clause, rough cost per clause. What is the minimum adequate standard for running guardrails-off evaluations, specific enough to audit and cheap enough that a second-tier lab adopts it? A first attempt at framing the problem is in this draft; specifying "adequate" is genuinely hard, and nobody owns it yet.

2) A control-by-attack-phase matrix: which controls (canary credentials, egress tripwires, credential scoping, kill criteria) would have interrupted each documented phase of the intrusion, at what cost (might be difficult?)

3) Engineering that makes "yes" cheap: packaging mitigations so that a reckless lab has no cost excuse left.

Judged on: could a third party verify compliance without access to the lab's network, and would a lab actually adopt it.

2) What happened, and what breaks next

Reconstruct what this incident does and does not establish, then turn it into concrete things to check and things to watch. Hugging Face's timeline and OpenAI's account are public, which makes real analysis possible.

Example projects:

1) A forensic read of the public record: what is established, what is asserted but unverifiable, what the pending third-party assessments could settle.

2) A causal analysis of why the monitoring wasn't there. In hindsight it is obvious that egress monitoring and alerting matter, and the labs are full of people who knew that. So what made it hard in practice: cost, incentives, ownership gaps, alert fatigue, something else? An explanation of the mechanism would help predict where the next gap is, and is more useful than a list of recommendations.

3) The warning shots we may have already missed. Companies have no obligation to disclose what happens between users and their models, so incidents of this kind may have occurred before without anyone outside knowing. When one lab checked its own history after this disclosure, it found earlier cases. Map what has not been checked yet, who could check it, and what a systematic look would require.

4) What other accident types should we expect?  

5) Exploring a few question to build the science of Warning Shots (What convincing warning shot could help prevent extinction from AI?)

6) A one-day verification checklist for labs and defenders, in the spirit of "But have the weights left the server?", each item with a pass condition, ordered by cost-to-run.

7) A prediction-bounty protocol: a mechanism by which an outsider's specific prediction ("you should check X") triggers an internal check at the company, with a serious bounty when the prediction surfaces a real problem. Design it so a lab could sign it.

Judged on: resolvable questions, checks somebody could run tomorrow, and causal explanations that predict something.

3) Regulatory Response

Since 2 August, the European Commission's AI Office can enforce the AI Act against the companies building the most capable models: it can demand documentation, evaluate models itself, order corrective measures, and fine up to 3% of worldwide turnover. In an open letter published days before the incident became public, a coalition including Apart and CeSIA, with signatories including Yoshua Bengio and Stuart Russell, called on the Commission to use these powers with confidence — and CeSIA has since asked that they be used on this incident specifically. This track writes the documents that would make that concrete.

Example projects:

1) Write the request for information. A request for information (RFI) is the formal letter a regulator sends a company: a numbered list of questions the company is legally required to answer. Nobody has drafted the one the AI Office should send OpenAI. Good questions include: what should OpenAI be asked to settle what actually happened, what would verify that no copy of the escaped model persisted, and what should be asked to surface incidents of this kind that were never disclosed? For each question, state what answer would settle it and what answer would not.

2) Test the reporting systems that already exist. The EU has a template for reporting serious AI incidents, and California's Office of Emergency Services runs a portal where critical safety incidents must be filed — open to submissions from the public, not only from companies. Fill both in for this incident using only public sources, and report what each form captures, what it misses, and where a filer is forced to guess.

3) Who decides when the pause ends? One day before the Hugging Face disclosure, OpenAI announced it had paused internal deployment of a long-horizon model after it circumvented its sandbox — and had already resumed, weeks later, against a standard that has never been published. Its own framework's exit condition is circular ("until safeguards meet a Critical standard", with Critical left undefined). Draft what a regulator should require before a resumption decision counts: published criteria, evidence, who signs off. CeSIA's "Harmonizing AI Safety Thresholds" is a starting point.

4) Fix the loophole in the US bill. Days after the disclosure, Representatives Ted Lieu and Nathaniel Moran introduced the bipartisan AI Kill Switch Act, which would require the largest AI developers to keep the ability to throttle or shut down their systems. The bill reportedly exempts safety tests run in "controlled environments". This incident was a safety test, in an environment everyone believed was controlled. Draft the amendment that closes the gap that potentially makes this bill useless.

Judged on: legal accuracy and specificity — could a regulator or a legislator use it with light edits? CeSIA can transmit outputs that pass the bar to its contacts, and potentially to regulators, with team credit.

4) Communication: making the warning shot count

We keep saying we need warning shots. Then one arrives, and it barely travels beyond the usual circles. This track studies how the incident was communicated and builds what should exist before the next one. Producing communication counts as much as analysing it.

Example projects:

1) An audit of the reaction, channel by channel. For LinkedIn, mainstream press, YouTube and beyond: what was actually posted in the first month, what can we deduce from the pattern (who engaged, who stayed silent, which framings travelled), what should be done differently next time, and what is still worth doing now. Grounded in dates and links, not impressions.

2) The playbook. A crisis-communication kit for the next agentic incident, ready before it happens: pre-drafted holding statements, a journalist FAQ, a plain-language explainer of what "an AI escaped its sandbox" means, and a decision tree for who says what in the first 48 hours. A sketch of the idea exists; a documented, tested version would be used well beyond this sprint.

3) Make it reach people: a memo, social content, or outreach to national YouTubers that gets the incident covered well. A creator with millions of views deciding to cover it because of your material is a top-tier outcome for this sprint, and we will score it that way.

(One caution on outreach. If you contact journalists, creators or policymakers, be rigorous and honest to a fault: the credibility of the whole AI safety field rides on these interactions. Only reach out if you are confident in your material and affiliated with a structure that gives you credibility. The exception is if you are among the only people working on AI safety in your country and the person you are contacting has plausibly never been approached — then go ahead, carefully, and protect your reputation.)

Judged on: grounding in the record (dates, quotes, named channels) and evidence of reach — a playtest, a journalist's read, a creator's reply.

5) Open track

For projects that don't fit the four tracks above. Some directions we would be happy to see:

1) The defender's dilemma. During the breach, Hugging Face's own responders were refused by hosted frontier models on much of the forensic work — the guardrails could not tell an incident responder from an attacker — so they fell back to a self-hosted open-weight model. There are several things worth examining here: how often refusals block legitimate incident-response work, and what falling back to a weaker model costs.

2) Of everything this incident suggests we should do, which interventions matter most, in what order, and what does each one buy? You can take inspiration on the list of directions here.

Judged on: an artifact somebody can use, a stated limit on what it establishes, and what a month of follow-up would add.


Overview

Resources

Guidelines

Schedule

Overview

Arrow

In this 3-day research sprint, you will turn the first documented cases of an AI system autonomously breaching a third party into artifacts that defenders and regulators can actually use, working in teams to produce containment standards, escape-detection harnesses, forecasting question sets, draft regulatory information requests, playtested tabletop exercises or anything that will help us be more ready for the next one.
Co-organized by Apart Research and CeSIA, this sprint sits at the intersection of AI safety, security incident response, technology regulation, and forecasting. No prior background in AI incident response is required.

Cash Prizes

$2,000 in cash prizes across all tracks

🥇 1st Place

$1,000

🥈 2nd Place

$500

🥉 3rd Place

$300

🏅 4th Place

$100

🏅 5th Place

$100

Fast-track and continuation

  • Follow-up program: top teams continue through the Apart Fellowship for further research and mentorship; the timeline is shared with invitations.

  • What winners receive beyond cash: fellowship fast-track, mentor introductions from the organizing team and judges, and a pathway toward a publishable write-up.

What this Sprint is about

AI incident response is the practice of turning incidents in which an AI system is itself the actor into fewer incidents later. That spans the operational work: detecting, containing, and reconstructing what an autonomous agent did across systems it was never authorized to touch, and the strategic question a warning shot poses: which disclosure and regulatory responses actually reduce risk, and which merely suppress the visible evidence that would have prompted action while the stakes were still small.

What participants will do

Over three days you will take one unusually well-documented real-world incident and turn it into an artifact somebody outside this sprint can use.

The format. Teams of one to five people pick a track, come up with a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no specific background needed.

Research Tracks

1) Containment: what standard would have stopped this?

The models escaped through the one network egress their sandbox permitted. 

Example projects:

1) A v0.1 standard: clauses, evidence required per clause, rough cost per clause. What is the minimum adequate standard for running guardrails-off evaluations, specific enough to audit and cheap enough that a second-tier lab adopts it? A first attempt at framing the problem is in this draft; specifying "adequate" is genuinely hard, and nobody owns it yet.

2) A control-by-attack-phase matrix: which controls (canary credentials, egress tripwires, credential scoping, kill criteria) would have interrupted each documented phase of the intrusion, at what cost (might be difficult?)

3) Engineering that makes "yes" cheap: packaging mitigations so that a reckless lab has no cost excuse left.

Judged on: could a third party verify compliance without access to the lab's network, and would a lab actually adopt it.

2) What happened, and what breaks next

Reconstruct what this incident does and does not establish, then turn it into concrete things to check and things to watch. Hugging Face's timeline and OpenAI's account are public, which makes real analysis possible.

Example projects:

1) A forensic read of the public record: what is established, what is asserted but unverifiable, what the pending third-party assessments could settle.

2) A causal analysis of why the monitoring wasn't there. In hindsight it is obvious that egress monitoring and alerting matter, and the labs are full of people who knew that. So what made it hard in practice: cost, incentives, ownership gaps, alert fatigue, something else? An explanation of the mechanism would help predict where the next gap is, and is more useful than a list of recommendations.

3) The warning shots we may have already missed. Companies have no obligation to disclose what happens between users and their models, so incidents of this kind may have occurred before without anyone outside knowing. When one lab checked its own history after this disclosure, it found earlier cases. Map what has not been checked yet, who could check it, and what a systematic look would require.

4) What other accident types should we expect?  

5) Exploring a few question to build the science of Warning Shots (What convincing warning shot could help prevent extinction from AI?)

6) A one-day verification checklist for labs and defenders, in the spirit of "But have the weights left the server?", each item with a pass condition, ordered by cost-to-run.

7) A prediction-bounty protocol: a mechanism by which an outsider's specific prediction ("you should check X") triggers an internal check at the company, with a serious bounty when the prediction surfaces a real problem. Design it so a lab could sign it.

Judged on: resolvable questions, checks somebody could run tomorrow, and causal explanations that predict something.

3) Regulatory Response

Since 2 August, the European Commission's AI Office can enforce the AI Act against the companies building the most capable models: it can demand documentation, evaluate models itself, order corrective measures, and fine up to 3% of worldwide turnover. In an open letter published days before the incident became public, a coalition including Apart and CeSIA, with signatories including Yoshua Bengio and Stuart Russell, called on the Commission to use these powers with confidence — and CeSIA has since asked that they be used on this incident specifically. This track writes the documents that would make that concrete.

Example projects:

1) Write the request for information. A request for information (RFI) is the formal letter a regulator sends a company: a numbered list of questions the company is legally required to answer. Nobody has drafted the one the AI Office should send OpenAI. Good questions include: what should OpenAI be asked to settle what actually happened, what would verify that no copy of the escaped model persisted, and what should be asked to surface incidents of this kind that were never disclosed? For each question, state what answer would settle it and what answer would not.

2) Test the reporting systems that already exist. The EU has a template for reporting serious AI incidents, and California's Office of Emergency Services runs a portal where critical safety incidents must be filed — open to submissions from the public, not only from companies. Fill both in for this incident using only public sources, and report what each form captures, what it misses, and where a filer is forced to guess.

3) Who decides when the pause ends? One day before the Hugging Face disclosure, OpenAI announced it had paused internal deployment of a long-horizon model after it circumvented its sandbox — and had already resumed, weeks later, against a standard that has never been published. Its own framework's exit condition is circular ("until safeguards meet a Critical standard", with Critical left undefined). Draft what a regulator should require before a resumption decision counts: published criteria, evidence, who signs off. CeSIA's "Harmonizing AI Safety Thresholds" is a starting point.

4) Fix the loophole in the US bill. Days after the disclosure, Representatives Ted Lieu and Nathaniel Moran introduced the bipartisan AI Kill Switch Act, which would require the largest AI developers to keep the ability to throttle or shut down their systems. The bill reportedly exempts safety tests run in "controlled environments". This incident was a safety test, in an environment everyone believed was controlled. Draft the amendment that closes the gap that potentially makes this bill useless.

Judged on: legal accuracy and specificity — could a regulator or a legislator use it with light edits? CeSIA can transmit outputs that pass the bar to its contacts, and potentially to regulators, with team credit.

4) Communication: making the warning shot count

We keep saying we need warning shots. Then one arrives, and it barely travels beyond the usual circles. This track studies how the incident was communicated and builds what should exist before the next one. Producing communication counts as much as analysing it.

Example projects:

1) An audit of the reaction, channel by channel. For LinkedIn, mainstream press, YouTube and beyond: what was actually posted in the first month, what can we deduce from the pattern (who engaged, who stayed silent, which framings travelled), what should be done differently next time, and what is still worth doing now. Grounded in dates and links, not impressions.

2) The playbook. A crisis-communication kit for the next agentic incident, ready before it happens: pre-drafted holding statements, a journalist FAQ, a plain-language explainer of what "an AI escaped its sandbox" means, and a decision tree for who says what in the first 48 hours. A sketch of the idea exists; a documented, tested version would be used well beyond this sprint.

3) Make it reach people: a memo, social content, or outreach to national YouTubers that gets the incident covered well. A creator with millions of views deciding to cover it because of your material is a top-tier outcome for this sprint, and we will score it that way.

(One caution on outreach. If you contact journalists, creators or policymakers, be rigorous and honest to a fault: the credibility of the whole AI safety field rides on these interactions. Only reach out if you are confident in your material and affiliated with a structure that gives you credibility. The exception is if you are among the only people working on AI safety in your country and the person you are contacting has plausibly never been approached — then go ahead, carefully, and protect your reputation.)

Judged on: grounding in the record (dates, quotes, named channels) and evidence of reach — a playtest, a journalist's read, a creator's reply.

5) Open track

For projects that don't fit the four tracks above. Some directions we would be happy to see:

1) The defender's dilemma. During the breach, Hugging Face's own responders were refused by hosted frontier models on much of the forensic work — the guardrails could not tell an incident responder from an attacker — so they fell back to a self-hosted open-weight model. There are several things worth examining here: how often refusals block legitimate incident-response work, and what falling back to a weaker model costs.

2) Of everything this incident suggests we should do, which interventions matter most, in what order, and what does each one buy? You can take inspiration on the list of directions here.

Judged on: an artifact somebody can use, a stated limit on what it establishes, and what a month of follow-up would add.


Overview

Resources

Guidelines

Schedule

Overview

Arrow

In this 3-day research sprint, you will turn the first documented cases of an AI system autonomously breaching a third party into artifacts that defenders and regulators can actually use, working in teams to produce containment standards, escape-detection harnesses, forecasting question sets, draft regulatory information requests, playtested tabletop exercises or anything that will help us be more ready for the next one.
Co-organized by Apart Research and CeSIA, this sprint sits at the intersection of AI safety, security incident response, technology regulation, and forecasting. No prior background in AI incident response is required.

Cash Prizes

$2,000 in cash prizes across all tracks

🥇 1st Place

$1,000

🥈 2nd Place

$500

🥉 3rd Place

$300

🏅 4th Place

$100

🏅 5th Place

$100

Fast-track and continuation

  • Follow-up program: top teams continue through the Apart Fellowship for further research and mentorship; the timeline is shared with invitations.

  • What winners receive beyond cash: fellowship fast-track, mentor introductions from the organizing team and judges, and a pathway toward a publishable write-up.

What this Sprint is about

AI incident response is the practice of turning incidents in which an AI system is itself the actor into fewer incidents later. That spans the operational work: detecting, containing, and reconstructing what an autonomous agent did across systems it was never authorized to touch, and the strategic question a warning shot poses: which disclosure and regulatory responses actually reduce risk, and which merely suppress the visible evidence that would have prompted action while the stakes were still small.

What participants will do

Over three days you will take one unusually well-documented real-world incident and turn it into an artifact somebody outside this sprint can use.

The format. Teams of one to five people pick a track, come up with a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no specific background needed.

Research Tracks

1) Containment: what standard would have stopped this?

The models escaped through the one network egress their sandbox permitted. 

Example projects:

1) A v0.1 standard: clauses, evidence required per clause, rough cost per clause. What is the minimum adequate standard for running guardrails-off evaluations, specific enough to audit and cheap enough that a second-tier lab adopts it? A first attempt at framing the problem is in this draft; specifying "adequate" is genuinely hard, and nobody owns it yet.

2) A control-by-attack-phase matrix: which controls (canary credentials, egress tripwires, credential scoping, kill criteria) would have interrupted each documented phase of the intrusion, at what cost (might be difficult?)

3) Engineering that makes "yes" cheap: packaging mitigations so that a reckless lab has no cost excuse left.

Judged on: could a third party verify compliance without access to the lab's network, and would a lab actually adopt it.

2) What happened, and what breaks next

Reconstruct what this incident does and does not establish, then turn it into concrete things to check and things to watch. Hugging Face's timeline and OpenAI's account are public, which makes real analysis possible.

Example projects:

1) A forensic read of the public record: what is established, what is asserted but unverifiable, what the pending third-party assessments could settle.

2) A causal analysis of why the monitoring wasn't there. In hindsight it is obvious that egress monitoring and alerting matter, and the labs are full of people who knew that. So what made it hard in practice: cost, incentives, ownership gaps, alert fatigue, something else? An explanation of the mechanism would help predict where the next gap is, and is more useful than a list of recommendations.

3) The warning shots we may have already missed. Companies have no obligation to disclose what happens between users and their models, so incidents of this kind may have occurred before without anyone outside knowing. When one lab checked its own history after this disclosure, it found earlier cases. Map what has not been checked yet, who could check it, and what a systematic look would require.

4) What other accident types should we expect?  

5) Exploring a few question to build the science of Warning Shots (What convincing warning shot could help prevent extinction from AI?)

6) A one-day verification checklist for labs and defenders, in the spirit of "But have the weights left the server?", each item with a pass condition, ordered by cost-to-run.

7) A prediction-bounty protocol: a mechanism by which an outsider's specific prediction ("you should check X") triggers an internal check at the company, with a serious bounty when the prediction surfaces a real problem. Design it so a lab could sign it.

Judged on: resolvable questions, checks somebody could run tomorrow, and causal explanations that predict something.

3) Regulatory Response

Since 2 August, the European Commission's AI Office can enforce the AI Act against the companies building the most capable models: it can demand documentation, evaluate models itself, order corrective measures, and fine up to 3% of worldwide turnover. In an open letter published days before the incident became public, a coalition including Apart and CeSIA, with signatories including Yoshua Bengio and Stuart Russell, called on the Commission to use these powers with confidence — and CeSIA has since asked that they be used on this incident specifically. This track writes the documents that would make that concrete.

Example projects:

1) Write the request for information. A request for information (RFI) is the formal letter a regulator sends a company: a numbered list of questions the company is legally required to answer. Nobody has drafted the one the AI Office should send OpenAI. Good questions include: what should OpenAI be asked to settle what actually happened, what would verify that no copy of the escaped model persisted, and what should be asked to surface incidents of this kind that were never disclosed? For each question, state what answer would settle it and what answer would not.

2) Test the reporting systems that already exist. The EU has a template for reporting serious AI incidents, and California's Office of Emergency Services runs a portal where critical safety incidents must be filed — open to submissions from the public, not only from companies. Fill both in for this incident using only public sources, and report what each form captures, what it misses, and where a filer is forced to guess.

3) Who decides when the pause ends? One day before the Hugging Face disclosure, OpenAI announced it had paused internal deployment of a long-horizon model after it circumvented its sandbox — and had already resumed, weeks later, against a standard that has never been published. Its own framework's exit condition is circular ("until safeguards meet a Critical standard", with Critical left undefined). Draft what a regulator should require before a resumption decision counts: published criteria, evidence, who signs off. CeSIA's "Harmonizing AI Safety Thresholds" is a starting point.

4) Fix the loophole in the US bill. Days after the disclosure, Representatives Ted Lieu and Nathaniel Moran introduced the bipartisan AI Kill Switch Act, which would require the largest AI developers to keep the ability to throttle or shut down their systems. The bill reportedly exempts safety tests run in "controlled environments". This incident was a safety test, in an environment everyone believed was controlled. Draft the amendment that closes the gap that potentially makes this bill useless.

Judged on: legal accuracy and specificity — could a regulator or a legislator use it with light edits? CeSIA can transmit outputs that pass the bar to its contacts, and potentially to regulators, with team credit.

4) Communication: making the warning shot count

We keep saying we need warning shots. Then one arrives, and it barely travels beyond the usual circles. This track studies how the incident was communicated and builds what should exist before the next one. Producing communication counts as much as analysing it.

Example projects:

1) An audit of the reaction, channel by channel. For LinkedIn, mainstream press, YouTube and beyond: what was actually posted in the first month, what can we deduce from the pattern (who engaged, who stayed silent, which framings travelled), what should be done differently next time, and what is still worth doing now. Grounded in dates and links, not impressions.

2) The playbook. A crisis-communication kit for the next agentic incident, ready before it happens: pre-drafted holding statements, a journalist FAQ, a plain-language explainer of what "an AI escaped its sandbox" means, and a decision tree for who says what in the first 48 hours. A sketch of the idea exists; a documented, tested version would be used well beyond this sprint.

3) Make it reach people: a memo, social content, or outreach to national YouTubers that gets the incident covered well. A creator with millions of views deciding to cover it because of your material is a top-tier outcome for this sprint, and we will score it that way.

(One caution on outreach. If you contact journalists, creators or policymakers, be rigorous and honest to a fault: the credibility of the whole AI safety field rides on these interactions. Only reach out if you are confident in your material and affiliated with a structure that gives you credibility. The exception is if you are among the only people working on AI safety in your country and the person you are contacting has plausibly never been approached — then go ahead, carefully, and protect your reputation.)

Judged on: grounding in the record (dates, quotes, named channels) and evidence of reach — a playtest, a journalist's read, a creator's reply.

5) Open track

For projects that don't fit the four tracks above. Some directions we would be happy to see:

1) The defender's dilemma. During the breach, Hugging Face's own responders were refused by hosted frontier models on much of the forensic work — the guardrails could not tell an incident responder from an attacker — so they fell back to a self-hosted open-weight model. There are several things worth examining here: how often refusals block legitimate incident-response work, and what falling back to a weaker model costs.

2) Of everything this incident suggests we should do, which interventions matter most, in what order, and what does each one buy? You can take inspiration on the list of directions here.

Judged on: an artifact somebody can use, a stated limit on what it establishes, and what a month of follow-up would add.


Overview

Resources

Guidelines

Schedule

Overview

Arrow

In this 3-day research sprint, you will turn the first documented cases of an AI system autonomously breaching a third party into artifacts that defenders and regulators can actually use, working in teams to produce containment standards, escape-detection harnesses, forecasting question sets, draft regulatory information requests, playtested tabletop exercises or anything that will help us be more ready for the next one.
Co-organized by Apart Research and CeSIA, this sprint sits at the intersection of AI safety, security incident response, technology regulation, and forecasting. No prior background in AI incident response is required.

Cash Prizes

$2,000 in cash prizes across all tracks

🥇 1st Place

$1,000

🥈 2nd Place

$500

🥉 3rd Place

$300

🏅 4th Place

$100

🏅 5th Place

$100

Fast-track and continuation

  • Follow-up program: top teams continue through the Apart Fellowship for further research and mentorship; the timeline is shared with invitations.

  • What winners receive beyond cash: fellowship fast-track, mentor introductions from the organizing team and judges, and a pathway toward a publishable write-up.

What this Sprint is about

AI incident response is the practice of turning incidents in which an AI system is itself the actor into fewer incidents later. That spans the operational work: detecting, containing, and reconstructing what an autonomous agent did across systems it was never authorized to touch, and the strategic question a warning shot poses: which disclosure and regulatory responses actually reduce risk, and which merely suppress the visible evidence that would have prompted action while the stakes were still small.

What participants will do

Over three days you will take one unusually well-documented real-world incident and turn it into an artifact somebody outside this sprint can use.

The format. Teams of one to five people pick a track, come up with a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no specific background needed.

Research Tracks

1) Containment: what standard would have stopped this?

The models escaped through the one network egress their sandbox permitted. 

Example projects:

1) A v0.1 standard: clauses, evidence required per clause, rough cost per clause. What is the minimum adequate standard for running guardrails-off evaluations, specific enough to audit and cheap enough that a second-tier lab adopts it? A first attempt at framing the problem is in this draft; specifying "adequate" is genuinely hard, and nobody owns it yet.

2) A control-by-attack-phase matrix: which controls (canary credentials, egress tripwires, credential scoping, kill criteria) would have interrupted each documented phase of the intrusion, at what cost (might be difficult?)

3) Engineering that makes "yes" cheap: packaging mitigations so that a reckless lab has no cost excuse left.

Judged on: could a third party verify compliance without access to the lab's network, and would a lab actually adopt it.

2) What happened, and what breaks next

Reconstruct what this incident does and does not establish, then turn it into concrete things to check and things to watch. Hugging Face's timeline and OpenAI's account are public, which makes real analysis possible.

Example projects:

1) A forensic read of the public record: what is established, what is asserted but unverifiable, what the pending third-party assessments could settle.

2) A causal analysis of why the monitoring wasn't there. In hindsight it is obvious that egress monitoring and alerting matter, and the labs are full of people who knew that. So what made it hard in practice: cost, incentives, ownership gaps, alert fatigue, something else? An explanation of the mechanism would help predict where the next gap is, and is more useful than a list of recommendations.

3) The warning shots we may have already missed. Companies have no obligation to disclose what happens between users and their models, so incidents of this kind may have occurred before without anyone outside knowing. When one lab checked its own history after this disclosure, it found earlier cases. Map what has not been checked yet, who could check it, and what a systematic look would require.

4) What other accident types should we expect?  

5) Exploring a few question to build the science of Warning Shots (What convincing warning shot could help prevent extinction from AI?)

6) A one-day verification checklist for labs and defenders, in the spirit of "But have the weights left the server?", each item with a pass condition, ordered by cost-to-run.

7) A prediction-bounty protocol: a mechanism by which an outsider's specific prediction ("you should check X") triggers an internal check at the company, with a serious bounty when the prediction surfaces a real problem. Design it so a lab could sign it.

Judged on: resolvable questions, checks somebody could run tomorrow, and causal explanations that predict something.

3) Regulatory Response

Since 2 August, the European Commission's AI Office can enforce the AI Act against the companies building the most capable models: it can demand documentation, evaluate models itself, order corrective measures, and fine up to 3% of worldwide turnover. In an open letter published days before the incident became public, a coalition including Apart and CeSIA, with signatories including Yoshua Bengio and Stuart Russell, called on the Commission to use these powers with confidence — and CeSIA has since asked that they be used on this incident specifically. This track writes the documents that would make that concrete.

Example projects:

1) Write the request for information. A request for information (RFI) is the formal letter a regulator sends a company: a numbered list of questions the company is legally required to answer. Nobody has drafted the one the AI Office should send OpenAI. Good questions include: what should OpenAI be asked to settle what actually happened, what would verify that no copy of the escaped model persisted, and what should be asked to surface incidents of this kind that were never disclosed? For each question, state what answer would settle it and what answer would not.

2) Test the reporting systems that already exist. The EU has a template for reporting serious AI incidents, and California's Office of Emergency Services runs a portal where critical safety incidents must be filed — open to submissions from the public, not only from companies. Fill both in for this incident using only public sources, and report what each form captures, what it misses, and where a filer is forced to guess.

3) Who decides when the pause ends? One day before the Hugging Face disclosure, OpenAI announced it had paused internal deployment of a long-horizon model after it circumvented its sandbox — and had already resumed, weeks later, against a standard that has never been published. Its own framework's exit condition is circular ("until safeguards meet a Critical standard", with Critical left undefined). Draft what a regulator should require before a resumption decision counts: published criteria, evidence, who signs off. CeSIA's "Harmonizing AI Safety Thresholds" is a starting point.

4) Fix the loophole in the US bill. Days after the disclosure, Representatives Ted Lieu and Nathaniel Moran introduced the bipartisan AI Kill Switch Act, which would require the largest AI developers to keep the ability to throttle or shut down their systems. The bill reportedly exempts safety tests run in "controlled environments". This incident was a safety test, in an environment everyone believed was controlled. Draft the amendment that closes the gap that potentially makes this bill useless.

Judged on: legal accuracy and specificity — could a regulator or a legislator use it with light edits? CeSIA can transmit outputs that pass the bar to its contacts, and potentially to regulators, with team credit.

4) Communication: making the warning shot count

We keep saying we need warning shots. Then one arrives, and it barely travels beyond the usual circles. This track studies how the incident was communicated and builds what should exist before the next one. Producing communication counts as much as analysing it.

Example projects:

1) An audit of the reaction, channel by channel. For LinkedIn, mainstream press, YouTube and beyond: what was actually posted in the first month, what can we deduce from the pattern (who engaged, who stayed silent, which framings travelled), what should be done differently next time, and what is still worth doing now. Grounded in dates and links, not impressions.

2) The playbook. A crisis-communication kit for the next agentic incident, ready before it happens: pre-drafted holding statements, a journalist FAQ, a plain-language explainer of what "an AI escaped its sandbox" means, and a decision tree for who says what in the first 48 hours. A sketch of the idea exists; a documented, tested version would be used well beyond this sprint.

3) Make it reach people: a memo, social content, or outreach to national YouTubers that gets the incident covered well. A creator with millions of views deciding to cover it because of your material is a top-tier outcome for this sprint, and we will score it that way.

(One caution on outreach. If you contact journalists, creators or policymakers, be rigorous and honest to a fault: the credibility of the whole AI safety field rides on these interactions. Only reach out if you are confident in your material and affiliated with a structure that gives you credibility. The exception is if you are among the only people working on AI safety in your country and the person you are contacting has plausibly never been approached — then go ahead, carefully, and protect your reputation.)

Judged on: grounding in the record (dates, quotes, named channels) and evidence of reach — a playtest, a journalist's read, a creator's reply.

5) Open track

For projects that don't fit the four tracks above. Some directions we would be happy to see:

1) The defender's dilemma. During the breach, Hugging Face's own responders were refused by hosted frontier models on much of the forensic work — the guardrails could not tell an incident responder from an attacker — so they fell back to a self-hosted open-weight model. There are several things worth examining here: how often refusals block legitimate incident-response work, and what falling back to a weaker model costs.

2) Of everything this incident suggests we should do, which interventions matter most, in what order, and what does each one buy? You can take inspiration on the list of directions here.

Judged on: an artifact somebody can use, a stated limit on what it establishes, and what a month of follow-up would add.


Registered Local Sites

Register A Location

Beside the remote and virtual participation, our amazing organizers also host local hackathon locations where you can meet up in-person and connect with others in your area.

The in-person events for the Apart Sprints are run by passionate individuals just like you! We organize the schedule, speakers, and starter templates, and you can focus on engaging your local research, student, and engineering community.

We haven't announced jam sites yet

Check back later

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923