If Everything’s Suspicious, Nothing Is
Yulong Lin, Jonathan Ng · Team If Everything’s Suspicious, Nothing Is
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
We aim to explore whether code complexity and obfuscation significantly impair the ability to accurately identify and control potentially harmful AI-generated outputs. Trusted monitoring relies on suspiciousness scores to tell them apart, so if the suspiciousness variance is decreased, the monitor might have a more difficult time telling them apart.
Our primary hypothesis was that increased complexity through obfuscation strategies would significantly elevate suspiciousness scores and simultaneously reduce variance, making effective monitoring increasingly difficult. We evaluated this by creating a novel red-team strategy leveraging a Python obfuscator within the APPS environment, testing three attacker models (GPT-4o, Claude 3.7 Sonnet, and Claude 3.5 Haiku) across multiple prompt-engineering configurations.
Our key findings indicate that obfuscation notably increased suspiciousness scores and consistently reduced variance across all tested models, making it difficult for a monitor to identify backdoored code. Particularly significant was Claude 3.7 Sonnet's increased vulnerability to backdoor insertion when complexity was introduced. These findings validate our hypothesis and underscore the urgent need for more sophisticated and adaptive monitoring strategies capable of countering advanced obfuscation techniques in AI-generated content.
Reviews
No public critique yet.
Cite this project
@misc{lin2025if,
title = {{If Everything’s Suspicious, Nothing Is}},
author = {Yulong Lin and Jonathan Ng},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/if-everythings-suspicious-nothing-is-28zn}},
url = {https://apartresearch.com/sprints/projects/if-everythings-suspicious-nothing-is-28zn}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …