LLM Agent Security: Jailbreaking Vulnerabilities and Mitigation Strategies
mohammed arsalan , Vishwesh bhat · team phoeniks
Submitted to Agent Security Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project investigates jailbreaking vulnerabilities in Large Language Model agents, analyzes their implications for agent security, and proposes mitigation strategies to build safer AI systems.

Reviews
It seems that most of the focus was in the literature review since these are all existing techniques, but the output is rather shallow and it is hard to know what is exactly the contribution. There could have been more emphasis in what is exactly the current state of jailbreaks, how your work advances our general knowledge of jailbreaks and how jailbreaks for agents are different from standard chatbot jailbreaks.
This submission talks about existing techniques but does not focus on how their work builds anything on top of them. The write up is very sparse.
This project provides a review of several methods to exploit vulnerabilities and as such jailbreak the LLM systems. The authors follow up with a discussion on potential societal and privacy implications for the same and discuss some mitigation strategies. The working demo shows one interesting example of prompt engineering that gets LLMs to answer with SQL injection attacks. Going forward I would love to see some potential mitigation strategies in the working demo
Cite this project
@misc{arsalan2024llm,
title = {{LLM Agent Security: Jailbreaking Vulnerabilities and Mitigation Strategies}},
author = {mohammed arsalan and Vishwesh bhat},
year = {2024},
month = oct,
note = {Submitted to Agent Security Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/llm-agent-security-jailbreaking-vulnerabilities-and-mitigation-strategies}},
url = {https://apartresearch.com/sprints/projects/llm-agent-security-jailbreaking-vulnerabilities-and-mitigation-strategies}
}More from Agent Security Hackathon
- 1st place by peer reviewView project: Diamonds are Not All You Need
Diamonds are Not All You Need
Diamonds are Not All You Need
This project tests an AI agent in a straightforward alignment problem. The agent is given creative freedom within a Minecraft world and is tasked with transforming a 100x100 radius of the world into diamond. It is …
- View project: Cross-model surveillance for emails handling
Cross-model surveillance for emails handling
Fluffy Vin
A system that implements cross-model security checks, where one AI agent (Agent A) interacts with another (Agent B) to ensure that potentially harmful actions are caught and mitigated before they can be executed. …
- View project: Inference-Time Agent Security
Inference-Time Agent Security
Inference-Time Agent Security
We take a first step towards automating model building for symbolic checking (eg formal verification, PDDL) of LLM systems.