SpecTrap: How Compliance Pressure Degrades AI-Generated Formal Specifications
Rahul Kumar · Team SpecTrap
Submitted to The Secure Program Synthesis Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
AI models are increasingly used to generate formal specifications and property-based tests for verification pipelines. We show that production-style prompts ("generate at least 8 properties, do not refuse") cause specification soundness to collapse: GPT-4o drops from 60% to 13% file-level correctness (p = 1.76×10⁻⁴, 252 generations, 3 models). The one-line remediation that fixes factual fabrication does not transfer to specification generation -- a novel negative finding. We release SpecTrap, a pip-installable tool that adversarially tests AI spec generators using Hypothesis validation and Z3 cross-checking, with all data and code open source.

Reviews
Specification synthesis typically focuses on two primary criteria: soundness and the ability to filter out incorrect programs. This work ensures soundness by combining the Hypothesis PBT library with the Z3 SMT solver. These tools complement one another to provide a robust verification framework. In contrast, specification strength is evaluated based on the diversity of generated specifications; results indicate that LLMs struggle in this area. Enhancing strength, potentially through iterative refinement, will be key to improve the quality of output.
Cite this project
@misc{kumar2026spectrap,
title = {{SpecTrap: How Compliance Pressure Degrades AI-Generated Formal Specifications}},
author = {Rahul Kumar},
year = {2026},
month = may,
note = {Submitted to The Secure Program Synthesis Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/spectrap-how-compliance-pressure-degrades-aigenerated-formal-specifications-lqbj}},
url = {https://apartresearch.com/sprints/projects/spectrap-how-compliance-pressure-degrades-aigenerated-formal-specifications-lqbj}
}More from The Secure Program Synthesis Hackathon
- View project: Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Lida Safety
Specifications for real systems do not exist as one-shot artifacts: the user's intent emerges as they discover edge cases, rewrite drafts, and react to failing tests. We present an iterative pipeline that takes this …
- View project: AgentSpecGap
AgentSpecGap
solo-team
This prototype extracts rules from system prompts, tool descriptions, and runtime config. Rules are classified into one of interface validation, authorization check, workflow ordering validation, runtime validation, …
- View project: SpecGap Arena
SpecGap Arena
Obligation Cartographers
SpecGap Arena is a benchmark and framework that exposes how incomplete specifications let plausible but incorrect code pass public tests. It synthesizes missing semantic obligations (security boundaries, invariants, …