Skip to content
Sprint projectJul 1, 2024

Detection of potentially deceptive attitudes using expression style analysis

Roland Pihlakas · Team ClarityNaut

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Detection of potentially deceptive attitudes using expression style analysis

Code (opens in new tab)
Share

My work on this hackathon consists of two parts: 1) As a sanity check, verifying the deception execution capability of GPT4. The conclusion is “definitely yes”. I provide a few arguments about when that is a useful functionality. 2) Experimenting with recognising potential deception by using an LLM-based text analysis algorithm to highlight certain manipulative expression styles sometimes present in the deceptive responses. For that task I pre-selected a small subset of input data consisting only of entries containing responses with elements of psychological influence. The results show that LLM-based text analysis is able to detect different manipulative styles in responses, or alternatively, attitudes leading to deception in case of internal thoughts.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{pihlakas2024detection,
  title = {{Detection of potentially deceptive attitudes using expression style analysis}},
  author = {Roland Pihlakas},
  year = {2024},
  month = jul,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/detection-of-potentially-deceptive-attitudes-using-expression-style-analysis}},
  url = {https://apartresearch.com/sprints/projects/detection-of-potentially-deceptive-attitudes-using-expression-style-analysis}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026