Sandbagging Detection via Consistency Probing
Atharshlakshmi Vijayakumar, Balakrishnan Vaisiya · Team hugginghands
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project detects sandbagging, when AI systems underperform in evaluative contexts while maintaining full capability in casual settings. We generate 50 paired prompts across seven reasoning domains and evaluate multiple LLMs under controlled conditions. By comparing performance on identical tasks framed as formal assessments versus casual interactions, and applying statistical analysis and a composite confidence score, we identify context-sensitive underperformance. The framework isolates incentive-driven behavior from task difficulty or ambiguity and can be integrated into red-teaming, auditing, and governance workflows to surface strategic underperformance before deployment.
Reviews
This project targets a known weakness in AI evaluation: models may behave differently in formal testing settings than in casual use, which can undermine safety assessments. The paired-prompt design and statistical analysis provide a clear and well-controlled way to surface such context-dependent performance drops, and testing across multiple models strengthens the evidence.
However, some observed differences may still reflect general prompt sensitivity rather than intentional sandbagging. Adding controls or framing variations specifically designed to separate these effects would help clarify the interpretation. Overall, this is a well-executed and promising evaluation approach that fits naturally into ongoing work on more reliable model assessment.
Cite this project
@misc{vijayakumar2026sandbagging,
title = {{Sandbagging Detection via Consistency Probing}},
author = {Atharshlakshmi Vijayakumar and Balakrishnan Vaisiya},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/sandbagging-detection-via-consistency-probing-32b5}},
url = {https://apartresearch.com/sprints/projects/sandbagging-detection-via-consistency-probing-32b5}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …