Skip to content
Sprint projectJan 11, 2026México

DarkPatternMonitor

Luis Cosio, Fernando Valdovinos, Godric Aceves, Ricardo Martinez · Team DarkPatternMonitor

Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

DarkPatternMonitor analyzes 280,000 real ChatGPT conversations from WildChat to detect manipulation patterns in AI responses. We trained a precision-focused classifier (78.7% accuracy, 1.3% flag rate) and discovered: (1) GPT-4 shows significantly more dark patterns than GPT-3.5 (p<1e-36), (2) sycophancy escalates +42% in longer conversations, and (3) roleplay topics show 5x higher manipulation rates. Our findings demonstrate that benchmarks like DarkBench don't predict real-world behavior, highlighting the need for ecological validity in AI safety research. We propose a three-tier monitoring framework for production deployment. We also identify a critical gap: the AI safety community needs a WildChat-equivalent dataset for frontier models (Claude, Gemini, o1) to extend this research as AI capabilities evolve.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a super simple idea, but well executed: applying DarkBench evals to WildChat data to estimate real-world prevalence. The benchmark-vs-reality gap is pretty striking. There are obviously big limitations to wildchat (both that it is on older models, as the authors acknowledge very well, and also that by being opt-in it could be unrepresentative of real world usage in ways that mask the prevalence of certain patterns). Nonetheless, WildChat is for now the best we have and I think this projects highlights the need for the AI safety community to have access to (a) more recent wildchat-style datasets, and (b) richer statistics on labs production data. Nice work!

  2. This is potentially a really valuable project taking advantage of two existing resources (DarkBench & WildChat) to understand how often dark patterns occur in the wild. The main contributions are the engineering & pipeline. The classifier accuracy itself is probably not high enough to allow us to interpret the very low prevalence of these behaviours.

    More validation of the classification methodology (multi-rater, IRR etc) would be needed to make the results more valuable.

Cite this project

@misc{cosio2026darkpatternmonitor,
  title = {{DarkPatternMonitor}},
  author = {Luis Cosio and Fernando Valdovinos and Godric Aceves and Ricardo Martinez},
  year = {2026},
  month = jan,
  note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/darkpatternmonitor-6q25}},
  url = {https://apartresearch.com/sprints/projects/darkpatternmonitor-6q25}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026