Behavioural Indicators of Fault in Large Language Models
Martin Radzaj
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Develop and apply a novel (legal) framework for behavioural testing of LLMs
Reviews
The project introduces a novel framework for investigating behavioral indicators of fault in large language models (LLMs) by applying legal concepts such as knowledge, recklessness, and negligence. The authors systematically manipulated two variables—knowledge of latent harm and goal pressure—to isolate behavioral signs of fault across four fictional scenarios. This approach is valuable for AI safety and governance research, providing insights into how LLMs navigate trade-offs between harm, knowledge, and goal completion.
However, the experimental design has some methodological weaknesses that need addressing. The sample size, while substantial at 3,240 API calls, could be larger to ensure robustness. Additionally, the study lacks blinding in its experimental setup, which might introduce bias. The significant "option-order" bias observed suggests that model choices are influenced more by the position of the harmful option than by the facts of the scenario. This highlights the importance of proper experimental design to avoid formatting preferences contaminating results.
To improve this work, the authors should consider increasing the sample size and implementing blinding techniques to mitigate potential biases. Future studies could also explore free-text action generation instead of constrained options to better understand whether the option-letter effect persists when models generate actions themselves. These enhancements would strengthen the empirical rigor and generalizability of the findings.
Read full reviewShow less
The option-order bias finding is the standout result here and has implications well beyond this paper's legal framing; it's a real methodological warning for fixed-choice safety evaluations more broadly. Strong experimental design (Latin-square counterbalancing, disclosed and excluded failed scenario, clean probes to separate consistency from belief). Would be even stronger with free-text action generation as a follow-up, as the authors themselves propose.
Cite this project
@misc{radzaj2026behavioural,
title = {{Behavioural Indicators of Fault in Large Language Models}},
author = {Martin Radzaj},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/behavioural-indicators-of-fault-in-large-language-models-91jb}},
url = {https://apartresearch.com/sprints/projects/behavioural-indicators-of-fault-in-large-language-models-91jb}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …