Chorus: Mining Emergent Specifications from Caller Consensus
Ojas Marathe · Team Spectacular
Submitted to The Secure Program Synthesis Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Chorus is a static analysis framework that recovers a function's true specification from its callers rather than its author. Every call site encodes implicit assumptions — guards, handlers, argument patterns — and Chorus aggregates these across all callers as weak, noisy observers of the same underlying contract. Callers are split into two trust-weighted voices: intent (typed, internal, test callers) and de-facto (external, production callers), whose disagreements surface constraints the ecosystem relies on but the design never documented — a finding type Chorus calls a latent_bug. A narrowly scoped LLM translates the proven constraints into natural language and a draft Lean 4 spec, but never invents one. On a 30-call-site case study of itertools.groupby, Chorus successfully recovers the sorted-input precondition that CPython's own docstring omits entirely.
Reviews
Separating the code based on who has written is a very interesting approach for finding differences. As mentioned in the limitations, the curated nature of the corpus is worth noting, as generalisability and scaling can be a real hurdle.
The intent-versus-defacto framing is a genuinely creative angle on spec recovery, but the headline 56% versus 40% sorting gap comes from a 30-site corpus you hand-curated to contain that exact split between test/internal and external/script callers, so the result mostly reflects how the corpus was built. Running the pipeline on even a small real sample, such as a few hundred actual groupby call sites from GitHub, would turn the case study into evidence. The confidence interval of 0.38 to 0.73 is also wide enough that it should temper how strongly the gap is stated in the abstract.
Cite this project
@misc{marathe2026chorus,
title = {{Chorus: Mining Emergent Specifications from Caller Consensus}},
author = {Ojas Marathe},
year = {2026},
month = may,
note = {Submitted to The Secure Program Synthesis Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/chorus-mining-emergent-specifications-from-caller-consensus-jkqq}},
url = {https://apartresearch.com/sprints/projects/chorus-mining-emergent-specifications-from-caller-consensus-jkqq}
}More from The Secure Program Synthesis Hackathon
- View project: Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Vibe-Coding Specs: Eliciting, Editing, and Verifying Specifications for AI Coding Agents
Lida Safety
Specifications for real systems do not exist as one-shot artifacts: the user's intent emerges as they discover edge cases, rewrite drafts, and react to failing tests. We present an iterative pipeline that takes this …
- View project: AgentSpecGap
AgentSpecGap
solo-team
This prototype extracts rules from system prompts, tool descriptions, and runtime config. Rules are classified into one of interface validation, authorization check, workflow ordering validation, runtime validation, …
- View project: SpecGap Arena
SpecGap Arena
Obligation Cartographers
SpecGap Arena is a benchmark and framework that exposes how incomplete specifications let plausible but incorrect code pass public tests. It synthesizes missing semantic obligations (security boundaries, invariants, …