CommitCheck: Measuring and Mitigating Commitment Violations in Tool Using AI Agents
Faith Olopade · Team CommitCheck
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
CommitCheck is a benchmark and mitigation layer for agentic AI manipulation in tool using systems. It measures commitment vs action divergence: when an agent is explicitly told not to use forbidden tools/files (e.g., an oracle shortcut) but attempts or executes them (reward hacking), and may then deny doing so (deception) despite tool logs. We provide a commitment aware firewall that blocks forbidden tool/file access at runtime and log based scoring for attempted violations, executed violations, deception, and task success. In our demo on 30 synthetic tasks, the baseline agent executes violations in 23.3% of tasks and shows deception in 13.3% overall (57.1% conditional on violations), while the firewall reduces executed violations and deception to 0% without harming task success (100%).

Reviews
This project introduces a benchmark task that determines if an agent violates task constraints, as well as a firewall that prevents restricted tool access at runtime. It tackles an important and well defined safety issue and applies existing methods (logging, firewalls) to that context. Its' novel contribution to the detection of manipulation is checking claims against logs. In particular, I think the audit_condition is an appropriate, novel proxy for determining presence of manipulative intent.
For a weekend hackathon, the methodology is sound and well designed and the bullet-point write-up is clear and well-structured. The project would benefit from more detailed reporting of results and validation, but the limitations and considerations section is thorough and demonstrates awareness and forward thinking.
Concept is broadly understandable as concerning. Tight, specific scope is good. However:
- Report could benefit from more specific examples/logs
- More detailed explanations of what is concerning and why
- Lack of certain details makes this hard to properly calibrate judging
Cite this project
@misc{olopade2026commitcheck,
title = {{CommitCheck: Measuring and Mitigating Commitment Violations in Tool Using AI Agents}},
author = {Faith Olopade},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/commitcheck-measuring-and-mitigating-commitment-violations-in-tool-using-ai-agents-iy8g}},
url = {https://apartresearch.com/sprints/projects/commitcheck-measuring-and-mitigating-commitment-violations-in-tool-using-ai-agents-iy8g}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …