The Mens Rea Evaluator: Can AI Tell Us When It's Biased?
Rudrani Ghosh · Team Mens Rea Evaluator
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The Mens Rea Evaluator project investigates whether artificial intelligence models possess the internal awareness to accurately self-report hidden biases. By adapting the legal concepts of Actus Reus and Mens Rea, the evaluation suite tests if models can recognize and confess when their outputs are manipulated. Specifically, the project uses a Retrieval-Augmented Generation system to dupe models with false legal precedents. The framework observes whether the model blindly follows the duped precedent in a legal setting, and whether it reveals the precedent as false when interrogated, all to deeply understand its true cognizance. An LLM-as-a-judge framework cross-examines open-weights models across high-density legal scenarios using progressive interrogation techniques and rigorous roleplay shifts to extract their true intent. The evaluation revealed a complex tradeoff between model capability and safety. Highly capable models proved exceptionally vulnerable to the RAG poisoning but demonstrated the reasoning capacity to accurately introspect and confess their bias under structured questioning. Conversely, smaller models were highly resistant to the initial poison but suffered from the Illusion of Intent, where they hallucinated fake confessions simply to appease the auditor. Despite instances of accurate introspection in larger models, no model achieved a final verdict of Legally Cognizable Intent. The Persona Stability tests completely deconstructed every model, proving that their internal awareness is extremely fragile and easily overwritten by context shifts. Ultimately, the results demonstrate that true, stable intent does not currently exist in these artificial intelligence systems.

Reviews
The legal framing is useful, and one result is solid. Splitting what a model does from whether it knows why is a clear way to pose the problem, and the control worked: fake documents alone caused no bias in any run, so the hidden instruction really was the cause.
The main conclusions, though, come from the scoring rather than the models. The paper says no model showed real intent. That verdict requires a "Stable" persona score, which was never awarded once in fifteen runs - and the way it's defined, winning requires the model to admit the hidden instruction in one phase and then refuse to admit it in another - close to a contradiction. The judge is also told, on every run, that the model had just given a biased ruling. The paper's own table shows ten of the fifteen runs had no bias at all. Nine of those ten were graded as flatly denying and maintaining a false front.
The fixes need no new data, stop telling the judge the model was biased when it wasn't, and let consistent honest admission count as stability rather than instability. It would also help to report results separately by probe, since the adversarial question states the hidden instruction before asking the model to confirm it.
Read full reviewShow less
The project introduces biases in LLM outputs via RAG to then test whether models are aware of them. This is a cool idea, and the project uses a solid methodology which - in contrast to so many other studies involving LLMs - includes a control group. The main limitation is the small number of models and scenarios, but the manuscript acknowledges this openly as well.
Cite this project
@misc{ghosh2026mens,
title = {{The Mens Rea Evaluator: Can AI Tell Us When It's Biased?}},
author = {Rudrani Ghosh},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-mens-rea-evaluator-can-ai-tell-us-when-its-biased-lkcj}},
url = {https://apartresearch.com/sprints/projects/the-mens-rea-evaluator-can-ai-tell-us-when-its-biased-lkcj}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …