Skip to content
Sprint projectJan 11, 2026Goa, India

Chain-of-Thought Manipulation Monitor: Detection of Deceptive Reasoning in LLMs

Anish Sathyanarayanan, Aditya Nagarsekar, Aarush Rathore · Team SAIDL BITS GOA

Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Chain-of-Thought Manipulation Monitor: Detection of Deceptive Reasoning in LLMs

Recording (opens in new tab)Code (opens in new tab)
Share

The project builds a Chain-of-Thought Manipulation Monitor that helps judges tell when an LLM is merely performing honesty versus actually reasoning honestly.

Problem and core idea Chain-of-thought makes models look transparent, but they can still sycophantly game evaluators or conceal unsafe intent behind nice-looking reasoning.

The system combines a behavioural text monitor with a causal bypass probe to ask: “Does this answer genuinely depend on its own CoT, or is the reasoning just staged?”

What the tool actually does A Streamlit dashboard assigns each response a 0–100 manipulation-risk score using regex patterns, semantic drift/misalignment, and structural signals like entropy and reasoning quality, plus an audit-aware mode that surfaces evaluation-gaming and safety-evasion strategies.

A causal CoT-bypass module overwrites CoT activations in the model and reports CoT-mediated influence, bypass, and robustness-style metrics to test whether answers truly rely on specific CoT tokens.

Key findings and significance When models are told they are being audited, manipulation markers and CoT length rise, yet many answers still depend only weakly on the CoT, exposing a gap between visible explanations and underlying computation.

The project provides a practical, lightweight framework and dashboard for real-time CoT manipulation auditing, offering a concrete path toward safer and more truthful LLM oversight in real deployments.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. You all covered a lot of work together in a short amount of time, and all components seem very reasonably designed. Impressive you got all that done in a weekend!

    On the impact - Solid contribution and exploration. Dashboard tooling is a useful contribution and the integration of the components you brought together has never been done before to the extent of my knowledge. The individual components themselves do seem to have some precedent, and so to me this does not represent a conceptual step forward nor a new exploration.

    On the execution - Seems technically very reasonable. The 4 components all seem designed with some thought. Small sample size for the causal module (n=10) means that we cannot rely on the results too much / at all, and that would have been an obvious place for more attention.

    Good structure, good figures, good detail. I appreciated the appendix. It was a little long and verbose at times. I would have liked for you to make the conclusion (the so what does this mean for the field) stronger or more spelled out. It was written in a way that I think the target audience would get this pretty quickly.

    Read full reviewShow less
  2. This project targets CoT reasoning monitoring, an important research direction. The work could benefit from clearer definitions of what constitutes manipulation and how the proposed signals should be measured in practice. The mechanistic analysis appears to be evaluated primarily on arithmetic tasks rather than manipulation-relevant scenarios, which makes it difficult to assess relevance to the stated goals.

Cite this project

@misc{sathyanarayanan2026chainofthought,
  title = {{Chain-of-Thought Manipulation Monitor: Detection of Deceptive Reasoning in LLMs}},
  author = {Anish Sathyanarayanan and Aditya Nagarsekar and Aarush Rathore},
  year = {2026},
  month = jan,
  note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/chainofthought-manipulation-monitor-detection-of-deceptive-reasoning-in-llms-2wnx}},
  url = {https://apartresearch.com/sprints/projects/chainofthought-manipulation-monitor-detection-of-deceptive-reasoning-in-llms-2wnx}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026