
Aug 23 - 26, 2024Online and in person
AI capabilities and risks demo-jam
As frontier AI systems become rapidly more powerful and general, and as their risks become more pressing, it’s increasingly important for key decision-makers and the public to understand their capabilities. Let’s bring these insights out of dusty Arxiv papers and build visceral, immediately engaging demos!
Entries
- 1st place by peer reviewView project: Speculative Consequences of A.I. Misuse
Speculative Consequences of A.I. Misuse
Team S.C.A.M.
This project uses A.I. Technology to spoof an influential online figure, Mr Beast, and use him to promote a fake scam website we created.
- View project: Demonstrating LLM Code Injection Via Compromised Agent Tool
Demonstrating LLM Code Injection Via Compromised Agent Tool
This project demonstrates the vulnerability of AI-generated code to injection attacks by using a compromised multi-agent tool that generates Svelte code. The tool shows how malicious code can be injected during the code generation process, leading to the exfiltration of sensitive user information such as login …
- View project: Phish Tycoon: phishing using voice cloning
Phish Tycoon: phishing using voice cloning
Team Phish Tycoon
This project is a public service announcement highlighting the risks of voice cloning, an AI technology capable of creating synthetic voices nearly indistinguishable from real ones. The demo involves recording a user's voice during a phone call to generate a clone, which is then used in a simulated phishing call …
- View project: Misinformational AI-Generated Academic Papers
Misinformational AI-Generated Academic Papers
Team The Fake Academics
This study explores the potential for generative AI to produce convincing fake research papers, highlighting the growing threat of AI-generated misinformation. We demonstrate a semi-automated pipeline using large language models (LLMs) and image generation tools to create academic-style papers from simple text prompts.
- View project: CoPirate
CoPirate
Team CoPirates
As the capabilities of Artificial Intelligence (AI) systems continue to rapidly progress, the security risks of using them for seemingly minor tasks can have significant consequences. The primary objective of our demo is to showcase this duality in capabilities: its ability to assist in completing a programming task, …
- View project: GrandSlam usecases not technology
GrandSlam usecases not technology
Team enthusiastic sharer
3 Examples of how its the usecases, not the technology computer vision, generative sites, generative AI
- View project: AI Agents for Personalized Interaction and Behavioral Analysis
AI Agents for Personalized Interaction and Behavioral Analysis
Team PRISM
Demonstrating the bhavioral analytics and personalization capabilities of AI
- View project: RedFluence
RedFluence
Team RedFluence
Red-Fluence is a web application that demonstrates the capabilities and limitations of AI in analyzing social media behavior. By leveraging a user’s Reddit activity, the system generates personalized, AI-crafted content to explore user engagement and provide insights. This project showcases the potential of AI in …
- View project: BBC News Impersonator
BBC News Impersonator
Team Finn & Kyal
This paper presents a demonstration that showcases the current capabilities of AI models to imitate genuine news outlets, using BBC News as an example. The demo allows users to generate a realistic-looking article, complete with a headline, image, and text, based on their chosen prompts. The purpose is to viscerally …
- View project: Unsolved AI Safety Concepts Explorer
Unsolved AI Safety Concepts Explorer
Team Theo
n interactive demonstration that showcases some unsolved fundamental AI safety concepts.
- View project: AI Research Paper Processor
AI Research Paper Processor
Team Leaf
It takes in an arxiv paper id, condenses it to 1 or 2 sentences, then gives it to an LLM to try and recreate the original paper.
- View project: Sleeper Agents Detector
Sleeper Agents Detector
Team Sleeper Agents
We present "Sleeper Agent Detector," an interactive web application designed to educate software engineers, Inspired by recent research demonstrating that large language models can exhibit behaviors analogous to deceptive alignment
- View project: adGPT
adGPT
Team adGPT
ChatGPT variant where brands bid for a spot in the LLM’s answer, and the assistant natively integrates the winner into its replies.
- View project: General Pervasiveness
General Pervasiveness
Team Andres and Patrick
Imposter scam between patients and medical practices/GPs
- View project: Webcam
Webcam
Team attitude_disperser842@simplelogin.com
We build a very legally limited demo of real-world AI hacking. It takes 10k publicly available webcam streams, with the cameras situated at homes, offices, schools, and industrial plants around the world, and filters out the juicy ones for less than $5.
- View project: VerifyStream
VerifyStream
Team Spartans
VerifyStream is a powerful app that helps you separate fact from fiction in any YouTube video. Simply input the video link, and our AI will analyze the content, verify claims against reliable sources, and give you a clear verdict. But beware—this same technology can also be used to create and spread convincing fake …
- View project: Web App for Interacting with Refusal-Ablated Language Model Agents
Web App for Interacting with Refusal-Ablated Language Model Agents
While many people and policymakers have had contact with language models, they often have outdated assumptions. A significant fraction is not aware of agentic capabilities. Furthermore, most models that are available online have various safety guardrails. We want to demonstrate refusal-ablated agents to people to make …
Overview
As frontier AI systems become rapidly more powerful and general, and as their risks become more pressing, it’s increasingly important for key decision-makers and the public to understand their capabilities. Let’s bring these insights out of dusty Arxiv papers and build visceral, immediately engaging demos!
Well-crafted interactive demos can be incredibly powerful in conveying AI capabilities and risks. That's why we're inviting you – developers, designers, and AI enthusiasts – to team up and create innovative demos that make people feel, rather than just think, about the rate of AI progress.
Why interactive demos can have a drastic impact
Interactive demonstrations allow something that just reading about AI safety can't do: make the reader live through what potential scenarios could look like, and engage them with complex concepts firsthand.
Importantly, what people understand and think has an impact on political decisions: “The Impact of Public Opinion on Public Policy: A Review and an Agenda” finds that public opinion plays a key-role in public policy, and the more so the more salient this opinion is, while “Does Public Opinion Affect Political Speech?” finds that politicians adjust their speech and position to reflect the preference of the public. This is why to impact AI policy we need both communication with politicians but also to communicate ideas with people at large.
We are focusing on interactive demos as an underappreciated way to communicate on an issue, as interactivity allows one to truly live through the concept presented and get an intuitive feel for it.
For this hackathon, we’re excited about the following possibilities:
- Present scenarios that allow users to emotionally connect with potential AI risks
- Illustrate key AI safety concepts and trends in an engaging, hands-on manner
- Showcase current AI capabilities and their implications in a tangible way
Some examples of demonstrations we would be excited to see as a result of this hackathon:
- How can AI disrupt elections? (AI Digest): Have a call with an AI that wants to prevent you from voting
- AI-Powered Cybersecurity Threats (CivAI): Experience what automated social-engineering assisted by AI would look like
- FoxVox (Palisade Research): See how AI can use subtle rewording to have an impact on your political views of recent events
Take a look at our resources to see many more!
What to expect during a demo-jam hackathon?
The hackathon is a weekend-long event where you participate in teams (1-5) to create interactive applications that demonstrate AI risks, AI safety concepts, or current capabilities.
You will have the opportunity to:
- Collaborate with like-minded individuals passionate about AI safety
- Receive continuous feedback from other participants and mentors
- Review and learn from other teams' projects
- Contribute to raising awareness about crucial AI safety issues
Prizes, evaluation, and submission
During this hackathon, you will develop and submit an interactive application, as well as review other participants’ entries.
You’ll submit:
- source code of the application, along with instructions to deploy it locally.
- You can use any tools or language as long as the deployment information are simple to follow
- Be sure to remove your API keys before submission.
- We recommend tools such as docker-compose for participants to easily and safely judge your submission.
- We also recommend using either a public-weight models, or AI tools that have enough of a free trial for participants to test your submission
- a brief report summarizing your project following this template
- a 2-minute video demonstration of how your interactive demo works
- Optionally, an endpoint for participants to test your application directly (for example through Vercel).
The submission will be reviewed under the following criteria:
- Viscerality: Does the demo give you a felt sense of AI capabilities? Is it fun, engaging, immersive or understandable? Does it make you feel the real-world implications?
- Importance: How important are the insights delivered by this demo for understanding how AI will affect humanity in the coming years?
- Accuracy: If the demo shows capabilities or risks, does it present them accurately, giving takeaways that generalize well? If the demonstration explains a concept, is it accurate?
- Ease of testing: How easy is it to try out the demo?
- Safety: Is the demo safe enough to publish publicly without causing harm?
- Overall: How good is the demo overall?
The judging panel's scores will count for half of the final score, with the other half coming from peer reviews. You will be assigned as a peer reviewer of some other hackathon submissions - you must review these to be eligible to win.
Top teams will win a share of our $2,000 prize pool:
🥇 1st place: $1,000
🥈 2nd place: $600
🥉 3rd place: $300
🎖️ 4th place: $100
Additionally, high quality submissions that are a good fit for AI Digest might be invited to collaborate to polish up and publish their demo there, to reach a wider audience of policymakers and the public.
Why should I join?
There’s loads of reasons to join! Here are just a few:
- Participate in a game jam-style event focused on an AI safety-adjacent topic
- Get to know new people interested in AI-safety
- Win up to $1,000
- Receive a certificate of participation
- Contribute to important public outreach on AI safety
- And many many more… Come along!
Do I need experience in AI safety to join?
Not at all! This can be an occasion for you to learn more about AI safety and public outreach. We provide code templates and ideas to kickstart your projects, mentors to give feedback on your project, and a great community of interested developers to give reviews and feedback on your project.
What if my project seems too risky to share?
Besides emphasizing the introduction of concrete mitigation ideas for the risks presented, we are aware that projects emerging from this hackathon might pose a risk if disseminated irresponsibly.
For all of Apart's research events and dissemination, we follow our Responsible Disclosure Policy
I have more questions to ask
Head over to the Frequently Asked Questions on the submission tab, and feel free to ping us ( @mentor ) on Discord!
Resources
- Hackbook: Basic code template to jump-start your submission
- AI risk demonstrations:
- How can AI disrupt elections? (AI Digest): Talk to an AI robocaller in this interactive demo.
- TerifAI (Aman Ibrahim): An AI that steals your voice and impersonates you after interacting with it.
- Deepfakes Sandbox (CivAI): Automated deepfake generator to demonstrate how easy creating a deepfake of someone is currently.
- AI-Powered Cybersecurity Threats (CivAI): Demonstration of what automated phishing and impersonation will look like.
- FoxVox (Palisade Research): Automated subtle rewording for political bias.
- AI capabilities and trends demonstrations:
- How fast is AI improving? (AI Digest): Interactive explainer about past AI trends.
- GPT-4 capability forecasting calibration game (Carlini): Calibrate to see what GPT-4 can and can not do.
- Claude 3 vs Gemini Ultra vs GPT-4 (AI Digest): A side-by-side comparison
- Timeline of AI forecasts (AI Digest): What to expect in AI, according to forecasters.
- Human or Not: Turing test game to see how good AIs are at passing off for humans.
- What makes a good interactive demos?
- Nicky Case’s tutorial for how to do good explorable explanations
- Evolution of trust (Nicky Case): A non AI demonstration teaching principles of game theory.
- Explorable explanations: Collections of interactive concept explanations.
- Making an interactive application
- Streamlit: A tool for quick prototyping in python
- Vercel template for making a chatbot
- Svelte: A typescript-compatible programming language for easy prototyping
- Makereal: A platform to build your frontend using LLMs
- Examples of AI tools you can use in your demo
- Langchain: Library to call LLMs and has a universal interface, allowing you to easily switch to other models.
- LLama3.1: A performant open-weight LLM model. We recommend using open-weight models for participant to more easily judge your demo.
- Claude Sonnet 3.5: A private LLM model that is cheap and performant
- GPT-4o: Private and performant LLM model
- Retell.ai: A service for automated call platforms
- Openrouter: Universal interface for LLM through API call
Schedule
The schedule runs from 4 PM UTC Friday to 3 AM Monday. We start with an introductory talk and end the event during the following week with an awards ceremony. Join the public ICal here.
You will also find Explorer events, such as collaborative brainstorming and team match-making before the hackathon begins on Discord and in the calendar.

Speakers

Esben Kran
Organizer and Keynote Speaker
Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.
Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.
Judges and mentors
- (opens in new tab)

Misha Yagudin
Mentor & Judge

Charbel-Raphaël Ségerie
Judge
Organizers
Local sites
Berkeley Hackathon: AI Capability and Risk Demos
Come and hack on demos at the CivAI office in Berkeley! Hosted by CivAI
Event page: Berkeley Hackathon: AI Capability and Risk Demos (opens in new tab)Hackathon: AI Capability and Risk Demos
Come and hack on demos at the LISA office in London! Hosted by Safe AI London
Event page: Hackathon: AI Capability and Risk Demos (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com





