Logit-Weighted Suspicion Scoring in AI Control
Maxime Cugnon de Sévricourt, Alexander Reinthal, Hamza Mooraj · Team Logit or Leave It
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
AI control relies on LLM monitors to detect suspicious behavior in untrusted agents. These monitors are typically run multiple times per evaluation to ensure statistically reliable suspicion scores, at significant cost in time and tokens. A natural way to reduce this cost is to extract more information from each monitor call by using the model's token probability distribution rather than its argmax output. We test this approach -- logit-weighted suspicion scoring, following \citet{liu-etal-2023-g} -- against standard point-estimate scoring in an AI control setting using 110 trajectories from AgentDojo, with both honest and attack agent policies. We find that logit-weighted scoring produces more continuous score distributions but does not improve AUC-ROC relative to point-estimate scoring. We interpret this as evidence that numerical anchoring in monitor outputs is rank-preserving in this setting, and that monitor capability rather than score granularity is the binding constraint on discrimination performance.
Reviews
I'm glad you published a null result and clearly said what dataset you used - agentdojo. Unfortunately, this research is not that useful because monitoring capability should have a significant increase if it has reasoning before it outputs a score. So, logit weighted scoring is not really relevant unless you split this into 2 turns (first turn giving reasoning, second turn giving answer). If that was done it would make this paper stronger.
Interesting idea, but the findings are not super groundbreaking.
Cite this project
@misc{sevricourt2026logitweighted,
title = {{Logit-Weighted Suspicion Scoring in AI Control}},
author = {Maxime Cugnon de Sévricourt and Alexander Reinthal and Hamza Mooraj},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/logitweighted-suspicion-scoring-in-ai-control-ktmf}},
url = {https://apartresearch.com/sprints/projects/logitweighted-suspicion-scoring-in-ai-control-ktmf}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …