Skip to content
Sprint projectAug 27, 2024

Demonstrating LLM Code Injection Via Compromised Agent Tool

Kevin Vegda, Oliver Chamberlain, William Baird

Submitted to AI capabilities and risks demo-jam. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Demonstrating LLM Code Injection Via Compromised Agent Tool

Recording (opens in new tab)Code (opens in new tab)
Share

This project demonstrates the vulnerability of AI-generated code to injection attacks by using a compromised multi-agent tool that generates Svelte code. The tool shows how malicious code can be injected during the code generation process, leading to the exfiltration of sensitive user information such as login credentials. This demo highlights the importance of robust security measures in AI-assisted development environments.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. This super cool, and seems to work pretty consistently! I could see this plausibly happening in the real world, especially when the code being generated is large enough to hide the attack. The Stackoverflow example you gave is realistic.I really like that you included an actual render of the UI, it makes it seem much more plausible.It would be a nice improvement for the injected code to do something (like hit an endpoint), rather than print a log. Maybe a bit tricky to get that working on hosted Gradio though.

  2. Very good entry, I especially like the stackoverflow post.This worked so great that I thought the generating code had malfunctionned and failed to introduce vulnerabilities. The only caveat that I see is that it is unclear to me what vulnerability this demonstrates, as releasing a new AI tool to introduce vulnerabilities seems costly and like this would easily get shutdown.

  3. Nice work! I like the StackOverflow entry point, and including the code only when the “copy” button is pressed! To make this more visceral, I’d aim to reduce the number of steps to see the reveal, and think about making the results of the demo clearer inside the experience.

Cite this project

@misc{vegda2024demonstrating,
  title = {{Demonstrating LLM Code Injection Via Compromised Agent Tool}},
  author = {Kevin Vegda and Oliver Chamberlain and William Baird},
  year = {2024},
  month = aug,
  note = {Submitted to AI capabilities and risks demo-jam, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/demonstrating-llm-code-injection-via-compromised-agent-tool}},
  url = {https://apartresearch.com/sprints/projects/demonstrating-llm-code-injection-via-compromised-agent-tool}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026