Skip to content
Sprint projectSep 14, 2026Marseille

Boundary Quorum Protocol: a verifiable authorization standard for agent evaluations

Lam Nguyen Le Tung · Team ThePenguin

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Boundary Quorum Protocol: a verifiable authorization standard for agent evaluations

Code (opens in new tab)
Share

In the July 2026 intrusion, a lab's evaluation let a model cross one trust boundary after another, and no single crossing had to be authorized by anyone. We propose the fix Track 1 asks for: the Boundary Quorum Protocol (BQP), a standard that sorts every boundary crossing into five risk tiers and sets the authorization each needs, from auto-execution, to an independent machine quorum, to a non-delegable human gate, to hard denial, with a hash-chained audit a third party can check. We calibrate the tiers on 44 machine-readable crossings from the public record, and we evaluate the standard against four authorization structures and four adversaries. Naive voting is not a containment control, while the full protocol executes none of the 32 unsafe crossings, at 22 human escalations, and we say where it still fails. A companion detection harness shows the gated crossings are catchable cheaply, with a rule that fires 66.5 hours before the internal-network pivot, and a companion attestation protocol, ECA, lets a third party verify a run from an exported bundle with no access to the lab's network.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The work here is credible, but not excellent. The post-hoc analysis of the single event and demonstration that the protocol works there is certainly some evidence, but seems to overfit to the narrow event in a post-hoc fashion. More consideration of the generalizability would be critical, and the reasoning for many of the asserted claims should be explained more clearly.

    It seems quite unclear that the proposed protocol is functional. The paper asserts that "The gated crossings can be caught cheaply and verifiably," but the expense of the detection of the predicates, the costs, and the accuracy are not assessed out of sample, or even in-sample. The list of events found is ad-hoc, and it's unclear that there is a way to automatically detect them other than parsing logs, which is expensive and of unclear accuracy. The list is useful, but far from complete, and partially automatically enforceable, but with an unclear false-positive rate.

    The effectiveness is also asserted, not argued for clearly. How does it know whether an identity or login is allowed or not? How does the verifier know if a read even has a side effect? (e.g. HTTP get requests are not actually read-only, and even when thy are, the remote server can take actions that create side effects.) How does the system reliably know if an transaction is high-impact, or prohibited?

    The writing was also substandard and messy LLM-generated content; asking a model to reread and rewrite section by section would have improved this, but nothing substitutes for humans carefully reading and ensuring that the work is clear, or better, actually writing or rewriting. (Using the acronym BQP was annoying on top of everything else.)

    Read full reviewShow less
  2. I was surprised to see that a naive multi-reviewer vote fails against collusion and outages.

    I'd assume it is the fix most teams would reach for first and your paper demonstrates it fails open under both.

    I would recommend replacing the scripted reviewers in the adversarial test with the real model from your live loop, since right now your independence claim is only proven against a hand-coded adversary.

  3. This paper addresses a foundational vulnerability in frontier model evaluations and autonomous agent deployments: the complete lack of independent authorization for high-risk boundary crossings during agent runs, as vividly illustrated by the July 2026 intrusion. The author introduces the Boundary Quorum Protocol (BQP) v0.1, a rigorous, risk-proportional authorization framework that gates agent tool execution using tiered risk levels (L0–L4), independence-constrained machine quorums, and non-delegable human gates. Calibrated against 44 documented events from the real incident, BQP is paired with a detection harness demonstrating early warning capabilities and the Evaluation Containment Attestation (ECA) v0.1 bundle for third-party verification. By proving that unsafe crossings can be intercepted with manageable operational overhead and verifiable audit trails, this work delivers a vital blueprint for auditable AI safety governance.

    Read full reviewShow less

Cite this project

@misc{tung2026boundary,
  title = {{Boundary Quorum Protocol: a verifiable authorization standard for agent evaluations}},
  author = {Lam Nguyen Le Tung},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/boundary-quorum-protocol-a-verifiable-authorization-standard-for-agent-evaluations-4hak}},
  url = {https://apartresearch.com/sprints/projects/boundary-quorum-protocol-a-verifiable-authorization-standard-for-agent-evaluations-4hak}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026