A Prospect Theoretic Approach to Agentic AI Safety
Lucas Irwin · Team Agentic Prosperity
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
In this project, I introduce Prospect-Theory CodeAct, a prompting framework that incorporates Kahneman & Tversky’s Prospect Theory into the decision loop of AI agents. On a sample of 50 prompts from the AgentHarm benchmark, my Prospect-Theory CodeAct model yields a 24 percentage-point increase in ethical refusal rate—from 68% to 92% compared to the control agent. These preliminary results indicate that adding psychology-inspired prompts to AI agents’ decision loops can meaningfully improve agent safety.
Reviews
Interesting, but is unclear how this helps accelerate defences against AI harms.
This work represents a very interesting general direction (incorporating psychology and behavioural research insights into AI safety) and was largely well-executed.
Unfortunately, the scores are capped by the lack of:
1. Ablations / baseline comparisons (e.g. just adding the consent part to the system prompt, without the rest of the prospect theory logic)
2. False positive analysis (a limitation you recognise) - it might be that this technique just shifts towards more refusals in general, including when refusals are not called for
3. Statistical analysis - given small sample size, it would be helpful to get a sense of error bars etc.
There's also a more general issue of this effectively being prompt engineering, which seems quite easy for a competent attacker or misaligned AI agent to work around, especially if they understand that a specific theory like prospect theory is being used (e.g. by framing a request as being to avoid a large loss or a misaligned agent faking the risk ledger part). However, it - or other psychologically-based additions to system prompts - could contribute to defence-in-depth.
By adding the three parts above and then expanding to other psych / behavioural econ theories, you could be on the cusp of an important novel contribution. And you did a pretty good execution for just a weekend!
Read full reviewShow less
Cite this project
@misc{irwin2025prospect,
title = {{A Prospect Theoretic Approach to Agentic AI Safety}},
author = {Lucas Irwin},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-prospect-theoretic-approach-to-agentic-ai-safety-ilil}},
url = {https://apartresearch.com/sprints/projects/a-prospect-theoretic-approach-to-agentic-ai-safety-ilil}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …