LLM-Assisted Ambiguity Detection in Regulatory Specifications
Kushagra Sharan · Team kshgrshrn
Submitted to The Secure Program Synthesis Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Regulatory software has a specification problem. Rules like "late filing attracts a penalty of Rs. 50 per day" are written for accountants and lawyers, not for engineers who need to know when the clock starts, whether there is a cap, and what happens when a supplier files late. The hard part is not implementing the rule once you understand it. The hard part is figuring out what the rule actually requires before you write any code. This project builds a lightweight pipeline that does one thing: reads a natural-language regulatory requirement and asks whether it is specific enough to implement. Three sequential LLM calls handle formalization, auditing, and scoring. The first rewrites the requirement as explicit bullet-point conditions. The second attacks the formalized version for edge cases and missing definitions. The third returns a structured JSON score and verdict. I ran it on seven GST-style compliance rules using Gemini. All seven came back underspecified, with repeated gaps around temporal boundaries, actor responsibilities, evidence requirements, and exception handling. None of that is surprising once you see it laid out, but having a system that surfaces those gaps before implementation starts is the point. The connection to secure synthesis is direct. A formally verified implementation of an incomplete specification is not safe, it is precisely wrong. The ambiguity finder sits one step before the formal methods pipeline and flags the questions that need answers first. The codebase is a single Python file with no external dependencies beyond the LLM API. A demo mode runs without any API key for reproducibility. The full Gemini output is committed under results/output.json.
Reviews
This is a good project exploration in an interesting space. Regulatory software is, indeed, an important use-case.
The security framing that a faithful implementation of an incomplete rule can let fraudulent claims pass is a good insight; the README is honest about scope and limits.
The main gap is in evaluation: all seven sampled requirements scored "underspecified" with no negative controls and no labeled ground truth, so the difference between <regulatory text is systematically underspecified> and <the prompt always says underspecified> is unclear.
A good next step is to add a handful of well-specified requirements as negative controls (so "adequate" can be returned) plus a few (human) expert-labeled cases to check that flagged ambiguities are real, since edge-case lists are a good deliverable.
Cite this project
@misc{sharan2026llmassisted,
title = {{LLM-Assisted Ambiguity Detection in Regulatory Specifications}},
author = {Kushagra Sharan},
year = {2026},
month = may,
note = {Submitted to The Secure Program Synthesis Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/llmassisted-ambiguity-detection-in-regulatory-specifications-k7v0}},
url = {https://apartresearch.com/sprints/projects/llmassisted-ambiguity-detection-in-regulatory-specifications-k7v0}
}More from The Secure Program Synthesis Hackathon
- View project: Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Lida Safety
Specifications for real systems do not exist as one-shot artifacts: the user's intent emerges as they discover edge cases, rewrite drafts, and react to failing tests. We present an iterative pipeline that takes this …
- View project: AgentSpecGap
AgentSpecGap
solo-team
This prototype extracts rules from system prompts, tool descriptions, and runtime config. Rules are classified into one of interface validation, authorization check, workflow ordering validation, runtime validation, …
- View project: SpecGap Arena
SpecGap Arena
Obligation Cartographers
SpecGap Arena is a benchmark and framework that exposes how incomplete specifications let plausible but incorrect code pass public tests. It synthesizes missing semantic obligations (security boundaries, invariants, …