SpecFault-Dafny: A Class-Stratified Scorecard for Verifier-Passing Specification Faults
Hugo Nguyen
Submitted to The Secure Program Synthesis Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A 45-item Dafny benchmark of verifier-passing specification faults: 30 intent/contract mismatches across six failure classes and nine domains, plus 15 clean controls. It scores five baseline validators (verifier-only, static escape-hatch scanning, symbolic checks, round-trip comparison, and a fixed-order hybrid) by coverage, clean-control false-positive rate, and per-class recall. The hybrid reaches 100% bad recall with a 0% observed clean-control false-positive rate, and every scored item ships with an inspectable evidence card.
Reviews
Practical and useful tool. If this could be run on some non curated tests, it would help get insight on what the hard problems are. Would also be interesting and more useful to add more testing tools to see more interactions.
This is a very well-scoped and well-implemented project. Great work!
Cite this project
@misc{nguyen2026specfaultdafny,
title = {{SpecFault-Dafny: A Class-Stratified Scorecard for Verifier-Passing Specification Faults}},
author = {Hugo Nguyen},
year = {2026},
month = may,
note = {Submitted to The Secure Program Synthesis Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/specfaultdafny-a-classstratified-scorecard-for-verifierpassing-specification-faults-q67o}},
url = {https://apartresearch.com/sprints/projects/specfaultdafny-a-classstratified-scorecard-for-verifierpassing-specification-faults-q67o}
}More from The Secure Program Synthesis Hackathon
- View project: Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Lida Safety
Specifications for real systems do not exist as one-shot artifacts: the user's intent emerges as they discover edge cases, rewrite drafts, and react to failing tests. We present an iterative pipeline that takes this …
- View project: AgentSpecGap
AgentSpecGap
solo-team
This prototype extracts rules from system prompts, tool descriptions, and runtime config. Rules are classified into one of interface validation, authorization check, workflow ordering validation, runtime validation, …
- View project: SpecGap Arena
SpecGap Arena
Obligation Cartographers
SpecGap Arena is a benchmark and framework that exposes how incomplete specifications let plausible but incorrect code pass public tests. It synthesizes missing semantic obligations (security boundaries, invariants, …