Behavioural Indicators of Fault in Large Language Models
Martin Radzaj
Develop and apply a novel (legal) framework for behavioural testing of LLMs
The project introduces a novel framework for investigating behavioral indicators of fault in large language models (LLMs) by applying legal concepts such as knowledge, recklessness, and negligence. The authors systematically manipulated two variables—knowledge of latent harm and goal pressure—to isolate behavioral signs of fault across four fictional scenarios. This approach is valuable for AI safety and governance research, providing insights into how LLMs navigate trade-offs between harm, knowledge, and goal completion.
However, the experimental design has some methodological weaknesses that need addressing. The sample size, while substantial at 3,240 API calls, could be larger to ensure robustness. Additionally, the study lacks blinding in its experimental setup, which might introduce bias. The significant "option-order" bias observed suggests that model choices are influenced more by the position of the harmful option than by the facts of the scenario. This highlights the importance of proper experimental design to avoid formatting preferences contaminating results.
To improve this work, the authors should consider increasing the sample size and implementing blinding techniques to mitigate potential biases. Future studies could also explore free-text action generation instead of constrained options to better understand whether the option-letter effect persists when models generate actions themselves. These enhancements would strengthen the empirical rigor and generalizability of the findings.
The option-order bias finding is the standout result here and has implications well beyond this paper's legal framing; it's a real methodological warning for fixed-choice safety evaluations more broadly. Strong experimental design (Latin-square counterbalancing, disclosed and excluded failed scenario, clean probes to separate consistency from belief). Would be even stronger with free-text action generation as a follow-up, as the authors themselves propose.
Cite this work
@misc {
title={
(HckPrj) Behavioural Indicators of Fault in Large Language Models
},
author={
Martin Radzaj
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


