Skip to content
Sprint projectMay 25, 2026Jersey City, NJ

AgentSpecGap

Amish · Team solo-team

Submitted to The Secure Program Synthesis Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

This prototype extracts rules from system prompts, tool descriptions, and runtime config. Rules are classified into one of interface validation, authorization check, workflow ordering validation, runtime validation, business-logic validation

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This work addresses the enforcement of policies on agent actions. It introduces a practical pipeline that extracts policies from data, compiles them into formal rules, and enforces them at runtime. The recovered specifications are then verified against both safe execution traces and mutated unsafe traces. However, because the current policy language is fixed and lacks expressiveness, integrating existing research on synthesizing minimally permissive permissions (i.e., strongest specifications) would be a direction for future work.

  2. AgentSpecGap is a strong and practical hackathon project. The core idea is very relevant: many LLM agents already contain safety rules scattered across prompts, tool descriptions, and runtime configs, but those rules are not enforced in a reliable or auditable way. This project turns that messy implicit policy layer into extracted candidate rules, compiles enforceable rules into a small policy IR, and checks tool calls at runtime before they execute.

    The best part of the project is that it is not just a concept. It has a runnable end-to-end pipeline with source artifacts, rule extraction, rule classification, policy compilation, runtime middleware, trace evaluation, and audit trails back to the original source span. That makes the project feel unusually concrete for a hackathon. The SQL and cloud-agent examples are also well chosen because they cover realistic failure cases: missing schema lookup, forbidden SQL actions, unapproved tables, missing authorization, public ACLs, unapproved buckets, and secret-labeled data being sent to public email.

    The enforcement design is sensible. Separating rules into enforce, review_only, and eval_required is a good engineering choice because not every English instruction should be turned into a hard runtime blocker. I also liked the check-category breakdown, since it makes clear which kinds of policies are currently supported and which ones still need semantic evaluation or richer runtime provenance.

    The main limitation is evaluation strength. The reported result of blocking 9/9 unsafe mutants with 0 false allows and 0 false blocks is encouraging, but the mutants are hand-authored and closely match the implemented operators. This is fine for a weekend prototype, but it does not yet show robustness to a broader space of realistic agent failures. A stronger version would automatically generate mutants from safe traces, test many more variants per rule, and include adversarial cases where the extracted rule is ambiguous or partially grounded.

    Another limitation is that the extraction path is not fully validated. The fixture fallback makes the demo deterministic, which is useful, but it means the strongest numbers do not fully measure the LLM extraction quality. The next important step is a labeled benchmark with human-grounded rules, comparing static-only extraction, LLM-only extraction, and the neurosymbolic extraction-plus-grounding pipeline. That would make the project much more convincing as a research contribution.

    Overall, this is one of the more practically useful projects in the hackathon setting. It identifies a real safety problem, builds a working enforcement pipeline, and is honest about what remains incomplete. The project would be even stronger with broader mutation coverage, real provenance propagation, and a stronger extraction benchmark, but the prototype already demonstrates a valuable direction.

    Read full reviewShow less

Cite this project

@misc{amish2026agentspecgap,
  title = {{AgentSpecGap}},
  author = {Amish},
  year = {2026},
  month = may,
  note = {Submitted to The Secure Program Synthesis Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/agentspecgap-bku6}},
  url = {https://apartresearch.com/sprints/projects/agentspecgap-bku6}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026