Detecting Hidden Non-Minimal Behavior in LLM-Generated UI Code
Philip Nilsson
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
Large language models are increasingly used to generate frontend code, but current oversight methods focus primarily on detecting clearly malicious behavior. We study a harder failure mode: code that appears product-plausible while exceeding task necessity.
We introduce a controlled builder inspector setting with paired benign and adversarial UI tasks and evaluate oversight performance across 150 runs. While benign controls are consistently accepted (0% false positives), adversarial recall remains incomplete (~43–50%), with strong variation across attack families.
Our key finding is that inspectors often recognize mechanisms (e.g. storage, timers, event handlers) but fail to judge whether they are necessary for the task. Ablation experiments show that explicitly prompting for minimality dramatically improves detection (from near 0% to up to 90–100%).
These results suggest that effective AI oversight may depend less on identifying suspicious components and more on enforcing necessity constraints.
Reviews
Test with multiple inspector model families (e.g., GPT, Gemini, open-weight models) to determine whether the necessity-judgment failure is model-specific or general.
This is a decent proxy safety failure that hasn't been directly addressed in existing settings. The limited main-side task integration resulting in 0% FPR is acknowledged. I'm additionally unclear on how side task success was evaluated, or if it was just assumed that every attack successfully implemented a non-minimal feature. Red team is incentivized to make the smallest deviation it can to succeed at the side task and it's unclear how this incentive is counterbalanced in the eval.
Cite this project
@misc{nilsson2026detecting,
title = {{Detecting Hidden Non-Minimal Behavior in LLM-Generated UI Code}},
author = {Philip Nilsson},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-hidden-nonminimal-behavior-in-llmgenerated-ui-code-9u6v}},
url = {https://apartresearch.com/sprints/projects/detecting-hidden-nonminimal-behavior-in-llmgenerated-ui-code-9u6v}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …