JARE: A Differential-Testing Workbench for Auditing AI-Generated Specifications
Allenna Tang, Jerry Chen, Elizabeth Kourbatski · Team JARE
Submitted to The Secure Program Synthesis Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
JARE is an auditing tool for AI-generated software specifications. When an AI writes both the rules a program should follow and the tests that check those rules, the tests can pass even when the rules are wrong. JARE finds concrete examples where an AI-generated specification disagrees with what the developer actually wanted, so a human can review and fix it.

Reviews
I'm afraid I can't spot a way that the proposed functionality is at all useful in program synthesis, vs. in carrying out evaluation of synthesis tools. The approach assumes collateral that would be very useful in the hands of an AI generating specs, yet the collateral is held out and only applied after generation.
One proposed usage mode is comparing a generated spec against a known-good spec. But if you have a known-good spec, why are you using AI to generate a spec from scratch, given just a natural-language description? Shouldn't the AI at least be shown the known-good spec, too?
Another proposed usage mode is evaluating generated specs against labeled test cases. Again, test cases are extremely useful to agentic AI-coding tools, so why aren't you sharing them?
Regardless of which scenario we focus on, the proposed functionality seems to follow standard practice in systematic testing, with no new conceptual contribution.
Read full reviewShow less
Promising experiments in a useful direction. One improvement could have been to compare different models in a multi-model setting: because eg ChatGPT and Claude are trained on different data, they can to some degree compensate for each others weaknesses, which may help significantly improve results in the no-gold deployment setting.
Cite this project
@misc{tang2026jare,
title = {{JARE: A Differential-Testing Workbench for Auditing AI-Generated Specifications}},
author = {Allenna Tang and Jerry Chen and Elizabeth Kourbatski},
year = {2026},
month = may,
note = {Submitted to The Secure Program Synthesis Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/jare-a-differentialtesting-workbench-for-auditing-aigenerated-specifications-o86j}},
url = {https://apartresearch.com/sprints/projects/jare-a-differentialtesting-workbench-for-auditing-aigenerated-specifications-o86j}
}More from The Secure Program Synthesis Hackathon
- View project: Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Lida Safety
Specifications for real systems do not exist as one-shot artifacts: the user's intent emerges as they discover edge cases, rewrite drafts, and react to failing tests. We present an iterative pipeline that takes this …
- View project: AgentSpecGap
AgentSpecGap
solo-team
This prototype extracts rules from system prompts, tool descriptions, and runtime config. Rules are classified into one of interface validation, authorization check, workflow ordering validation, runtime validation, …
- View project: SpecGap Arena
SpecGap Arena
Obligation Cartographers
SpecGap Arena is a benchmark and framework that exposes how incomplete specifications let plausible but incorrect code pass public tests. It synthesizes missing semantic obligations (security boundaries, invariants, …