Intent Inspector - Protecting Against Prompt Injections for Agent Tool Misuse
Oliver Morris, Gerard Boxo Corominas · Team Intent Inspector
Submitted to Agent Security Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
AI agents are powerful because they can affect the world via tool calls. This is a target for bad actors. We present protection against prompt injection aimed at tool calls in agents.
Reviews
This project addresses an extremely important and highly impactful problem. Corrupted tool calls can indeed cause havoc with autonomous agents. The approach is promising, though I would have liked to see more extensive experimentation with the Intent Inspector. Despite this, it's a solid hackathon project that tackles a crucial issue in AI safety.
Interesting project protecting against prompt injection attacks. Its focus on analyzing the intent behind actions makes it a highly intelligent and effective solution. The project’s ability to provide an additional layer of security is impressive, and the detailed approach taken makes it a standout contribution to agent safety.
• Good idea and implementation trying to develop a prompt injection detector. I particularly liked the direct comparison with Lakera’s product and found surprising that your implementation runs smoother than Lakera’s off-the-shelf product.
• It shows that there is room for more competitors in the for-profit AIS space and that one can get pretty far with in very short time. However, regarding safety it does not really move the field forward.
This project tackles the timely issue of prompt injection attacks in LLM-based agents using an "Intent Inspector" approach. I found the use of ASB data and the comparison with Lakera intriguing, especially the surprising performance of the smaller LLM model. The authors acknowledge limitations with attack primitiveness and temperature sensitivity which are valuable areas for future exploration.
Cite this project
@misc{morris2024intent,
title = {{Intent Inspector - Protecting Against Prompt Injections for Agent Tool Misuse}},
author = {Oliver Morris and Gerard Boxo Corominas},
year = {2024},
month = oct,
note = {Submitted to Agent Security Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/intent-inspector-protecting-against-prompt-injections-for-agent-tool-misuse}},
url = {https://apartresearch.com/sprints/projects/intent-inspector-protecting-against-prompt-injections-for-agent-tool-misuse}
}More from Agent Security Hackathon
- 1st place by peer reviewView project: Diamonds are Not All You Need
Diamonds are Not All You Need
Diamonds are Not All You Need
This project tests an AI agent in a straightforward alignment problem. The agent is given creative freedom within a Minecraft world and is tasked with transforming a 100x100 radius of the world into diamond. It is …
- View project: Cross-model surveillance for emails handling
Cross-model surveillance for emails handling
Fluffy Vin
A system that implements cross-model security checks, where one AI agent (Agent A) interacts with another (Agent B) to ensure that potentially harmful actions are caught and mitigated before they can be executed. …
- View project: Inference-Time Agent Security
Inference-Time Agent Security
Inference-Time Agent Security
We take a first step towards automating model building for symbolic checking (eg formal verification, PDDL) of LLM systems.