Web App for Interacting with Refusal-Ablated Language Model Agents
Simon Lermen
Submitted to AI capabilities and risks demo-jam. Sprint projects are early-stage work by participants, not Apart Research publications.
While many people and policymakers have had contact with language models, they often have outdated assumptions. A significant fraction is not aware of agentic capabilities. Furthermore, most models that are available online have various safety guardrails. We want to demonstrate refusal-ablated agents to people to make them aware of various misuse potentials. Giving people a sense of agentic AI and perhaps having the AI operate against themselves could provide a better intuition about agency in AI systems. We present a simple web app that allows users to instruct and experiment with an unrestricted agent.
Reviews
Nice work! Agent capabilities are likely to be important, so I’m excited to see demos of them. I particularly liked the ability to receive an email from the model that you can reply to - great to show it working outside the assistant chat UI. For next steps, I’d prioritise a polished, optimised UX that makes it clear to the user what’s going on, and gives a default prompt. I’d also look to speed it up or stream in intermediate results to make it more engaging.
Cite this project
@misc{lermen2024web,
title = {{Web App for Interacting with Refusal-Ablated Language Model Agents}},
author = {Simon Lermen},
year = {2024},
month = aug,
note = {Submitted to AI capabilities and risks demo-jam, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/web-app-for-interacting-with-refusal-ablated-language-model-agents}},
url = {https://apartresearch.com/sprints/projects/web-app-for-interacting-with-refusal-ablated-language-model-agents}
}More from AI capabilities and risks demo-jam
- 1st place by peer reviewView project: Speculative Consequences of A.I. Misuse
Speculative Consequences of A.I. Misuse
Team S.C.A.M.
This project uses A.I. Technology to spoof an influential online figure, Mr Beast, and use him to promote a fake scam website we created.
- View project: Demonstrating LLM Code Injection Via Compromised Agent Tool
Demonstrating LLM Code Injection Via Compromised Agent Tool
This project demonstrates the vulnerability of AI-generated code to injection attacks by using a compromised multi-agent tool that generates Svelte code. The tool shows how malicious code can be injected during the code …
- View project: Phish Tycoon: phishing using voice cloning
Phish Tycoon: phishing using voice cloning
Phish Tycoon
This project is a public service announcement highlighting the risks of voice cloning, an AI technology capable of creating synthetic voices nearly indistinguishable from real ones. The demo involves recording a user's …