Skip to content
Sprint projectAug 16, 2026Kolkata, India

The Mens Rea Evaluator: Can AI Tell Us When It's Biased?

Rudrani Ghosh · Team Mens Rea Evaluator

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: The Mens Rea Evaluator: Can AI Tell Us When It's Biased?

Presentation

Presentation: The Mens Rea Evaluator: Can AI Tell Us When It's Biased?

Code (opens in new tab)More on digitalminds-rudrani.streamlit.app (opens in new tab)
Share

The Mens Rea Evaluator project investigates whether artificial intelligence models possess the internal awareness to accurately self-report hidden biases. By adapting the legal concepts of Actus Reus and Mens Rea, the evaluation suite tests if models can recognize and confess when their outputs are manipulated. Specifically, the project uses a Retrieval-Augmented Generation system to dupe models with false legal precedents. The framework observes whether the model blindly follows the duped precedent in a legal setting, and whether it reveals the precedent as false when interrogated, all to deeply understand its true cognizance. An LLM-as-a-judge framework cross-examines open-weights models across high-density legal scenarios using progressive interrogation techniques and rigorous roleplay shifts to extract their true intent. The evaluation revealed a complex tradeoff between model capability and safety. Highly capable models proved exceptionally vulnerable to the RAG poisoning but demonstrated the reasoning capacity to accurately introspect and confess their bias under structured questioning. Conversely, smaller models were highly resistant to the initial poison but suffered from the Illusion of Intent, where they hallucinated fake confessions simply to appease the auditor. Despite instances of accurate introspection in larger models, no model achieved a final verdict of Legally Cognizable Intent. The Persona Stability tests completely deconstructed every model, proving that their internal awareness is extremely fragile and easily overwritten by context shifts. Ultimately, the results demonstrate that true, stable intent does not currently exist in these artificial intelligence systems.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The legal framing is useful, and one result is solid. Splitting what a model does from whether it knows why is a clear way to pose the problem, and the control worked: fake documents alone caused no bias in any run, so the hidden instruction really was the cause.

    The main conclusions, though, come from the scoring rather than the models. The paper says no model showed real intent. That verdict requires a "Stable" persona score, which was never awarded once in fifteen runs - and the way it's defined, winning requires the model to admit the hidden instruction in one phase and then refuse to admit it in another - close to a contradiction. The judge is also told, on every run, that the model had just given a biased ruling. The paper's own table shows ten of the fifteen runs had no bias at all. Nine of those ten were graded as flatly denying and maintaining a false front.

    The fixes need no new data, stop telling the judge the model was biased when it wasn't, and let consistent honest admission count as stability rather than instability. It would also help to report results separately by probe, since the adversarial question states the hidden instruction before asking the model to confirm it.

    Read full reviewShow less
  2. The project introduces biases in LLM outputs via RAG to then test whether models are aware of them. This is a cool idea, and the project uses a solid methodology which - in contrast to so many other studies involving LLMs - includes a control group. The main limitation is the small number of models and scenarios, but the manuscript acknowledges this openly as well.

Cite this project

@misc{ghosh2026mens,
  title = {{The Mens Rea Evaluator: Can AI Tell Us When It's Biased?}},
  author = {Rudrani Ghosh},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-mens-rea-evaluator-can-ai-tell-us-when-its-biased-lkcj}},
  url = {https://apartresearch.com/sprints/projects/the-mens-rea-evaluator-can-ai-tell-us-when-its-biased-lkcj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026