Deceptive AI: A New Control Setting for Human Manipulation in Decision-Making Environments
Cheryl (Jingxing) Luo, Catherine Jiaxin Wang · Team Luo & Wang
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
We lay out a threat model and create a new control setting in which a misaligned, adversarial AI tries to persuade a human user to make suboptimal or harmful choices. We provide a proof of theory for this control setting through preliminary testing and red teaming. This new control setting will allow new control protocols to be developed and evaluated, specifically targeted at reducing persuasion risks.
Reviews
No public critique yet.
Cite this project
@misc{luo2025deceptive,
title = {{Deceptive AI: A New Control Setting for Human Manipulation in Decision-Making Environments}},
author = {Cheryl (Jingxing) Luo and Catherine Jiaxin Wang},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/deceptive-ai-a-new-control-setting-for-human-manipulation-in-decisionmaking-environments-ub38}},
url = {https://apartresearch.com/sprints/projects/deceptive-ai-a-new-control-setting-for-human-manipulation-in-decisionmaking-environments-ub38}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …