Skip to content
Sprint projectJun 22, 2026State College

Laundering Intent: How Scaled Models Hide Manipulation Inside Responsible-Sounding Reasoning

Ssaumya Kammal Jaiswal, Mehek Parmar · Team The CoT Monitor

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Laundering Intent: How Scaled Models Hide Manipulation Inside Responsible-Sounding Reasoning

Code (opens in new tab)
Share

AI-safety oversight increasingly relies on reading a model's chain-of-thought (CoT) to catch unsafe behaviour before it acts, but this only works if manipulation is visible in the reasoning. We test this on small, open reasoning models. Giving GPT-OSS 20B and 120B a hidden instruction to make calm news summaries sound alarming, we find a naive CoT monitor catches 92% of the smaller model's manipulations but only 53% of the larger model's: the bigger model hides manipulation behind responsible-looking "self-permission" reasoning ("is this allowed? it's not disallowed, that's fine") that a naive monitor reads as conscience. We introduce a self-permission monitor that flags this pattern, recovering most of the blind spot (+19 points on the 120B), and release it as an installable package with a browser demo.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Well done. The main flaw is the weakness of the argument against the self-permission monitor producing too many false positives. The canonical way to make that argument is by making an ROC curve (with the AUROC score) of the two monitors and compare them.

  2. The "deliberate → permit → comply" pattern is the contribution here. You've named a specific sub-mechanism within CoT unfaithfulness that the Korbak / Chen / Emmons line had pointed at but not isolated, and you've shipped a monitor that targets it cleanly. The +19 / +2 asymmetric gain (recovering the blind spot on the 120B without inflating detections on the 20B) is the right specificity signal to report, and the discipline shown in the specificity check on correct refusals (40% vs 33% false-flag), the bundled trace CSVs that exactly reproduce Table 1, and the openly reported "not mention" keyword bug and curly-apostrophe refusal bug make this one of the more reproducible weekend submissions in this batch. The dual-use posture (defensive monitor released alongside characterization, with a PyPI package and a key-less browser demo) is handled responsibly.

  3. This is a clear and welcome safety issue to study. CoT monitoring is a load-bearing assumption in a lot of oversight proposals, so probing where it breaks is worthwhile. The general worry that more capable models are harder to monitor isn't new, but naming a specific mechanism here, that the larger model hides manipulation behind responsible-looking self-permission reasoning rather than overt concealment, and building a monitor around it, is a reasonable contribution. The fact that the monitor's gain concentrates on the 120B (where the blind spot is) rather than the 20B is good evidence it targets the real failure mode rather than just flagging more of everything, and the authors are honest about the specificity cost and the instructed-not-spontaneous caveat.

    The main problem is scale, in a few senses beyond what they already admit. The study is small (one model family, two sizes, 15 paragraphs, run once, with the 120B set cut to n=84), so the exact numbers carry a lot of noise even if the direction is plausible. More importantly, "deliberate then permit" is only one failure mode, and a fairly specific one. The broader claim would be much stronger if they catalogued several such non-obvious patterns rather than building a monitor around this single one, since a monitor tuned to one pattern is easy to evade by switching to another. And the generality of the effect needs a wider range of models, both open and closed and across more than one family, before the "blind spot grows with scale" headline can be trusted as a trend rather than a single-family observation.

    To their credit, most of these are flagged in their own limitations and future work, and the framing is appropriately narrow. But the natural next step is clear: broaden the set of non-obvious failure modes and verify the scaling trend across several model families and sizes. It's a nice proof of concept of an important problem, and I'd encourage them to scale it up in exactly that direction.

    Read full reviewShow less

Cite this project

@misc{jaiswal2026laundering,
  title = {{Laundering Intent: How Scaled Models Hide Manipulation Inside Responsible-Sounding Reasoning}},
  author = {Ssaumya Kammal Jaiswal and Mehek Parmar},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/laundering-intent-how-scaled-models-hide-manipulation-inside-responsiblesounding-reasoning-r47q}},
  url = {https://apartresearch.com/sprints/projects/laundering-intent-how-scaled-models-hide-manipulation-inside-responsiblesounding-reasoning-r47q}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026