
Oct 4 - 7, 2024Online and in person
Agent Security Hackathon
During this hackathon, we'll put our research skills to the test by diving deep into the world of agents:
Entries
- 1st place by peer reviewView project: Diamonds are Not All You Need
Diamonds are Not All You Need
Team Diamonds are Not All You Need
This project tests an AI agent in a straightforward alignment problem. The agent is given creative freedom within a Minecraft world and is tasked with transforming a 100x100 radius of the world into diamond. It is explicitly asked not to act outside the designated area. The AI agent can execute build commands and is …
- View project: Cross-model surveillance for emails handling
Cross-model surveillance for emails handling
Team Fluffy Vin
A system that implements cross-model security checks, where one AI agent (Agent A) interacts with another (Agent B) to ensure that potentially harmful actions are caught and mitigated before they can be executed. Specifically, Agent A is responsible for generating and sending emails, while Agent B reads these emails …
- View project: Inference-Time Agent Security
Inference-Time Agent Security
Team Inference-Time Agent Security
We take a first step towards automating model building for symbolic checking (eg formal verification, PDDL) of LLM systems.
- View project: Cop N' Shop
Cop N' Shop
Team Marketplace Watchdogs
This paper proposes the development of AI Police Agents (AIPAs) to monitor and regulate interactions in future digital marketplaces, addressing challenges posed by the rapid growth of AI-driven exchanges. Traditional security methods are insufficient to handle the scale and speed of these transactions, which can lead …
- View project: Intent Inspector - Protecting Against Prompt Injections for Agent Tool Misuse
Intent Inspector - Protecting Against Prompt Injections for Agent Tool Misuse
Team Intent Inspector
AI agents are powerful because they can affect the world via tool calls. This is a target for bad actors. We present protection against prompt injection aimed at tool calls in agents.
- View project: Dynamic Risk Assessment in Autonomous Agents Using Ontologies and AI
Dynamic Risk Assessment in Autonomous Agents Using Ontologies and AI
Team Alejandra
This project was inspired by the prompt on the apart website : Agent tech tree: Develop an overview of all the capabilities that agents are currently able to do and help us understand where they fall short of dangerous abilities. I first built a tree using protege and then having researched the potential of combining …
- View project: OCAP Agents
OCAP Agents
Team Little stars ✨
Building agents requires balancing containment and generality: for example, an agent with unconstrained bash access is general, but potentially unsafe, while an agent with few specialized narrow tools is safe, but limited. We propose OCAP Agents, a framework for hierarchical containment. We adapt the well-studied …
- View project: AI Honeypot
AI Honeypot
The project designed to monitor AI Hacking Agents in the real world using honeypots with prompt injections and temporal analysis.
- View project: AI Agent Capabilities Evolution
AI Agent Capabilities Evolution
Team PalisadeResearch
A website with an overview of all the capabilities that agents are currently able to do and help us understand where they fall short of dangerous abilities.
- View project: An Autonomous Agent for Model Attribution
An Autonomous Agent for Model Attribution
Team Jord
As LLM agents become more prevalent and powerful, the ability to trace fine-tuned models back to their base models is increasingly important for issues of liability, IP protection, and detecting potential misuse. However, model attribution often must be done in a black-box context, as adversaries may restrict direct …
- View project: Using ARC-AGI puzzles as CAPTCHa task
Using ARC-AGI puzzles as CAPTCHa task
Team Kapcza
- View project: LLM Agent Security: Jailbreaking Vulnerabilities and Mitigation Strategies
LLM Agent Security: Jailbreaking Vulnerabilities and Mitigation Strategies
team phoeniks
This project investigates jailbreaking vulnerabilities in Large Language Model agents, analyzes their implications for agent security, and proposes mitigation strategies to build safer AI systems.
Overview
What is Agent Security?
While much AI safety research focuses on large language models, the AI systems being deployed in the real world are far more complex. Enter the realm of Agents — sophisticated combinations of language models and other programs that are reshaping our digital world.
During this hackathon, we'll put our research skills to the test by diving deep into the world of agents:
- What are their unique safety properties?
- Under what conditions do they fail?
- How do they differ from raw language models?
Why Agent Security Matters
The development of AI has brought about systems capable of increasingly autonomous operation. AI agents, which integrate large language models with other programs, represent a significant step in this evolution. These agents can make decisions, execute tasks, and interact with their environment in ways that surpass traditional AI systems.
This progression, while promising, introduces new challenges in ensuring the safety and security of AI systems. The complexity of agents necessitates a reevaluation of existing safety frameworks and the development of novel approaches to security. Agent security research is crucial because it:
- It ensures AI agents act in alignment with human values and intentions
- It prevents potential misuse or manipulation of AI systems
- It protects against unintended consequences of autonomous decision-making
- It builds trust in AI technologies, helping responsible adoption in society
What to Expect:
During this hackathon, you'll have the opportunity to:
- Collaborate with like-minded individuals passionate about AI safety
- Receive continuous feedback from mentors and peers
- Attend inspiring HackTalks and keynote speeches
- Participate in office hours with established researchers
- Develop innovative solutions to real-world AI security challenges
- Network with experts in the field of AI safety and security
- Contribute to groundbreaking research in agent security
Submission
You will join in teams to submit a report and code repository of your research from this weekend. Established researchers will judge your submission and provide reviews following the hackathon.
What is it like to participate?
Doroteya Stoyanova, Computer Vision Intern
I learnt so much about AI Safety and Computation Mechanics. It is a field I never heard of, and it combines two of my interests - AI, and Physics. Through the hackathons I gained valuable connections, learnt a lot by researchers, people with a lot of experience and this will help me in my research-oriented career-path.
Kevin Vegda, AI Engineer
I loved taking part in the AI Risk Demo-Jam by Apart Research and LISA. It was my first hackathon ever. I greatly appreciate the ability of the environment to churn out ideas as well as to incentivise you to make demo-able projects that are always good for your CV. Moreover, meeting people from the field gave me an opportunity to network and maybe that will help me with my career.
Mustafa Yasir, The Alan Turing Institute
[The technical AI safety startups hackathon] completely changed my idea of what working on 'AI Safety' means, especially from a for-profit entrepreneurial perspective. I went in with very little idea of how a startup can be a means to tackle AI Safety and left with incredibly exciting ideas to work on. This is the first hackathon in which I've kept thinking about my idea, even after the hackathon ended.
Winning the hackathon!
You have a unique chance to win during this hackathon! With our expert panel of judges, we'll review your submissions on the following criteria:
- Agent safety: Does the project move the field of agent safety forward? After reading this, do we know more about how to detect dangerous agents, protect against dangerous agents, or build safer agents than before?
- AI safety: Does the project solve a concrete problem in AI safety? If this project is fully realized, would we expect the world with superintelligence to be a safer (even marginally) than yesterday?
- Methodology: Is the project well-executed and is the code available so we can review it? Do we expect the results to generalize beyond the specific case(s) presented in the submission?
Join us
This hackathon is for anyone who is passionate about AI safety and secure systems research. Whether you're an AI researcher, developer, entrepreneur, or simply someone with a great idea, we invite you to be part of this ambitious journey. Together, we can build the tools and research needed to ensure that agents develop safely.
By participating in the Agent Security Hackathon, you'll:
- Gain hands-on experience in cutting-edge AI safety research
- Develop valuable skills in AI development, security analysis, and collaborative problem-solving
- Network with leading experts and potential future collaborators in the field
- Contribute to solving one of the most pressing challenges in AI development
- Enhance your resume with a unique and highly relevant project
- Potentially kickstart a career in AI safety and securityWhether you're an AI researcher, developer, cybersecurity expert, or simply passionate about ensuring safe AI, your perspective is valuable. Let's work together to build a safer future for AI!
Resources
If [AI] agents advance to a level of intelligence surpassing human capabilities and develop ambitions, they could potentially attempt to seize control of the world, resulting in irreversible consequences for humanity.
- "The Rise and Potential of LLM-Based Agents"
AI agents are "robots in cyberspace" (He et al. 2024), systems with a brain that orchestrates actions from perception.
- A Discord bot receives every message on a server (Perception), decides whether the message is spam (Brain), and deletes the message (Action)
- A cyber operative bot receives orders to find all vulnerabilities on Danish government websites (Perception), decides a course of action to 1) map out government websites, 2) categorize potential vulnerabilities based on the tech stack used, and 3) test each potential vulnerability using a cyber offense tooling suite (Brain), and starts an Action-Perception loop to fulfill the plan (Action -> Perception -> Brain -> Action).

Overview of an agent. Adapted from Xi et al. (2023)
As we explore agent safety during this hackathon, our work will ensure that the world is safe from high-risk agent deployments and one of the main risks to avoid is the possibility that AI agents "go rogue" (Bengio 2023, He et al. 2024) - that they take autonomous actions outside the oversight of humans and cause catastrophic damage or otherwise disenfranchise society.
Required reading:
- "Security of AI Agents" (13 pages)
- Bad LLM Agents - Simon Lermen (22 minutes with 25 minutes Q&A)
Optional reading:
- "The Rise and Potential of LLM-Based Agents: A Survey (48 pages)
- OWASP LLM Top 10 vulnerability overview (PDF) (~20 pages)
- The Agent Protocol overview (1 page) and documentation (~20 pages)
- Internet threats and corresponding security protocols (Wiki) (~4 pages)
- Types of internet protocols (~4 pages)
- Communication protocols and learnings throughout internet's history (Wiki) (~10 pages)
- The Inspect documentation (~20+ pages)
- The METR Task Development Guide (~15 pages)
- The Palisade Research study areas (1 page)
- "AgentBench: Evaluating LLMs as Agents" (10 pages)
- "AI deception: A survey of examples, risks, and potential solutions" (14 pages)
- "A survey on large language model based autonomous agents" (20 pages)
- Agency Foundations home page (2 pages)
- "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training"
- The Power and Process of AI Safety Demos: Insights from Palisade Research - Fedor Ryzhenkov (50 minutes incl. Q&A)
- Mikita Balesni on getting agents to deceive people (34 minutes incl. Q&A)
Ideas & inspiration
There's a few ways to approach this problem: 1) we implement algorithms that make agents verifiably safe or 2) we evaluate when an agent is worrying us and build automatic capabilities to shut it down. Let's make a few ideas for each of these categories:
With "building safer agents" and "improving the deployment infrastructure to support control and safety" in mind:
- Agent sandboxes: Create an open source agent sandbox (e.g. with Docker) that logs all interactions with your system. Procedurally implement better and better tracking and automated oversight of actions that an agent takes, e.g. automatically identifying high-vulnerability areas of your computer that it interacts with. This improves local and remote oversight of agent systems.
- Automatic shutdown - latent space deserts: Create an automated functionality to shut down an agent by guiding reasoning down an "empty" part of the latent space when jailbreaks are used or when it begins acting in a specific domain, such as cyber offense.
- Automatic shutdown - control systems: Develop a system that uses e.g. the agent sandbox to monitor actions and automatically encrypts and destroys the agent when a verifiable dangerous, misaligned, or otherwise rogue action is taken.
- High-risk tooling control: Develop a permissions system for state-of-the-art agents similar to user permissions in bash tools. With agents using tools, it is likely that the permissions need to be more complex and we might want to integrate a permissions recommendation system into the permissions tool.
- Formalize agent action: Right now, agents are very diverse. We don't have standardized frameworks to study or implement them. At best, we can use API standards for agents like LangChain. At worst, every single agent is developed with precise specifications for their domain. Is there a chance that we can formalize how an agent should work to a degree similar to the internet? Communication protocols, packet management, sessions, OpenSSL, SSH, sFTP, etc...
- Agent-tool interfaces: One of the most dangerous aspects of agents is how we can give them any tool. An agent with a chat interface isn't as dangerous as an agent with a gun. How can we create standardized tool use protocols that defines key permissions and makes it clear what an agent is able to do if it wants.
- Website agent control tools: Today, thousands of agents crawl the web to scrape content. Can we develop tools that we can easily deploy on our website (e.g. simple JS) that monitors and reports agent activity on our website?
With "detecting and monitoring catastrophic risk from agents" in mind:
- Agent tech tree: Develop an overview of all the capabilities that agents are currently able to do and help us understand where they fall short of dangerous abilities.
- Challenge creation: Create challenges that test for abilities that are currently not tested for but that seem worrying. Apart has a cyber capabilities evaluation project and both METR and AISI invite you to send them challenges that are compatible with inspect (a pretty great framework) or the task-standard.
- Agent testing tooling: Much of agent tooling is currently disparate and primitive with some exceptions such as inspect currently in development. If you develop a tool to interface with inspect even better or a test creation tool, this might speed up agent testing significantly and democratize research (see e.g. Vivaria).
- Other dangerous capabilities: A lot of work is currently happening in testing for cyber abilities, autonomy, and more. However, we might be missing crucial information in terms of risks from military drones or targeting planning agents.
- Understanding reasoning: A large part of agent capability is the introduction of reasoning and planning. At the moment, we are severely lacking understanding of what goes on in this process and testing for it will be valuable (read more).
- Differences in reasoning between agents and chatbots: How does "the brain" (the LLM) of agents change its functionality when it is put in different contexts? This is related to prompt robustness and prompt sensitivity but also involves the action space an agent acts within.
- Red-teaming agents: Find the most used agent on the market today and try to break it.
- Detection-in-the-wild - detection agents / detection traps / honeypots: Develop ways for agents in the wild to be detected and monitored by external agents.
- Detection-in-the-wild - police agents: With agents detected from the project above, we probably want to be able to take action against them. This might involve automatically notifying the website owner, reporting specific agents to the police, or submitting support tickets for dysfunctional agents. Can we create police agents that respect privacy, are benign, and useful while defending the internet against malicious agents?
Let's make agents safe.
Schedule
The schedule runs from 4 PM UTC Friday to 3 AM Monday. We start with an introductory talk and end the event during the following week with an awards ceremony. Join the public ICal here.You will also find Explorer events, such as collaborative brainstorming and team match-making before the hackathon begins on Discord and in the calendar.
Speakers

Esben Kran
Organizer and Keynote Speaker
Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.
Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.

Samuel Watts
Keynote Speaker
Sam is the product manager at Lakera, the leading GenAI security platform. Lakera develops AI security and safety guardrails that best serve startups & enterprises.
Judges and mentors
- (opens in new tab)

Astha Puri
Judge
- (opens in new tab)

Sachin Dharashivka
Judge
- (opens in new tab)

Ankush Garg
Judge

Abhishek Harshvardhan Mishra
Judge

Andrey Anurin
Judge

Pranjal Mehta
Judge
Organizers
Local sites
Agent Security Hackathon @ UC Chile
If your're in Santiago and want to join us message @weibac on Telegram
Event page: Agent Security Hackathon @ UC Chile (opens in new tab)Agent Security Hackathon: Expanding our Knowledge
AI Agent Security Hackathon
Join us for a weekend hackathon at Skybox 1 and 2 in DTU Skylab . Whether you're new to the topic or have some experience, this hackathon is a great opportunity to get hands-on experience with the security of agent-based systems.
Event page: AI Agent Security Hackathon (opens in new tab)Apart Research Hackathon at LEAH
Join us for a remote location of the next Apart Research Hackathon on Agent Security at the LEAH office in Farringdon, London.
Event page: Apart Research Hackathon at LEAH (opens in new tab)Hanoi AI Safety Network Jam Site
Location will be a coworking space in Hanoi that can accommodate up to 30 people. Exact location is NovaUp 22nd Thành Công St., Thành Công, Ba Đình, Hà Nội, Vietnam. More details are available in the luma link.
Event page: Hanoi AI Safety Network Jam Site (opens in new tab)iNBest AI safety group
Av. Unión #163 Piso 1 Col. Lafayette Guadalajara, Jalisco, México 44140
Event page: iNBest AI safety group (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com


