Exploration Chat-Based Social Engineering
Ivan Lee, Amber Chang · Team Dizzy Cloud
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
With the advancement of AI Agent technologies, language models have increasingly demonstrated human-like characteristics, particularly in applications involving companionship and psychological counseling. As these models become more proficient in simulating human conversation, new social engineering attack strategies have emerged in the domain of fraud. Malicious actors can now exploit large language models (LLMs) in conjunction with publicly available user information to engage in highly personalized dialogue. Once a sufficient level of familiarity is established, these interactions may lead to phishing attempts or the extraction of sensitive personal data. This study proposes a method for investigating social engineering attacks driven by language models, referred to as ECSE (Exploring Chat-based Social Engineering). We utilize several open-source models—GPT-4o, GPT-4o-mini, LLaMA 3.1, and DeepSeek-V3—as the foundation for this framework. Through prompt engineering techniques, we collect experimental data in a sandbox to evaluate the conversational capability and operational efficiency of these models within a static social context.

Reviews
No public critique yet.
Cite this project
@misc{lee2025exploration,
title = {{Exploration Chat-Based Social Engineering}},
author = {Ivan Lee and Amber Chang},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/exploration-chatbased-social-engineering-fkqa}},
url = {https://apartresearch.com/sprints/projects/exploration-chatbased-social-engineering-fkqa}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …