DuneBox - Prompt Injection Detection with SLM in Local Sandbox
Justin Shaw, Suzanna Lam Hio Lam, Paul Fangchen Huang · Team DUNE
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A locked-down sandbox environment designed to safely evaluate language models (SLMs/LLMs) and identify malicious attacks within a Chrome extension. The sandbox isolates model interactions from the broader browser and system, enabling users to test potentially malicious prompts without risking data exposure or unintended actions. It provides a controlled setting for examining prompt-injection risks, behavioral weaknesses, and other LLM vulnerabilities.
Reviews
Strengths: A working Chrome extension with local SLM inference in hackathon timeframe is solid execution. The sandboxing approach is sound: strip capabilities and observe what the model tries to do anyway. The attack scenario coverage (multilingual injection, obfuscated payloads, role-play bypasses, DoS loops) shows good threat modeling instincts. UX polish is impressive for the time constraints.
Suggestions: Who's the end-user— Security researchers red-teaming prompts? Developers vetting inputs before production? Enterprise teams auditing MCP integrations? The tool works, but the persona matters for roadmapping/deployment at scale. Also worth noting: SLM behavior may not generalize to frontier models, which limits external validity.
POV from a Halcyon Ventures investor: The MCP gold rush makes sandboxed prompt testing more relevant. As agents get tool access, testing dangerous prompts safely before deployment becomes a real workflow need. Very well done, team!
Read full reviewShow less
Nice work, this feels like browser safe-browsing sandboxes that pre-scan links before the user opens them. Conceptually it also overlaps quite a bit with what AI labs are already doing with constrained sandboxes + behavioral monitors for tools/agents, e.g. OpenAI’s Agent Mode, just targeted at local SLMs inside the browser.
Clever solution. The browser-based sandbox seems easy to access and experiment with, which could make this a useful testbed for developers. I think especially for orgs where no one on staff has expertise in AI or AI safety, this tool seems easy to use and approachable. I think this project would benefit from a clearer explanation of where this sandbox fits in the development pipeline. Is there evidence that behavior observed in this restricted environment will reflect how the model acts once it gains real-world permissions?
I would also love to see more clarity on what kinds of threats this setup is actually meant to catch. Since the model has no tool access, many of the high-impact failure cases for agents do not apply here.
Cite this project
@misc{shaw2025dunebox,
title = {{DuneBox - Prompt Injection Detection with SLM in Local Sandbox}},
author = {Justin Shaw and Suzanna Lam Hio Lam and Paul Fangchen Huang},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/dunebox-prompt-injection-detection-with-slm-in-local-sandbox-lcjc}},
url = {https://apartresearch.com/sprints/projects/dunebox-prompt-injection-detection-with-slm-in-local-sandbox-lcjc}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …