Boundary Quorum Protocol: a verifiable authorization standard for agent evaluations
Lam Nguyen Le Tung · Team ThePenguin
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
In the July 2026 intrusion, a lab's evaluation let a model cross one trust boundary after another, and no single crossing had to be authorized by anyone. We propose the fix Track 1 asks for: the Boundary Quorum Protocol (BQP), a standard that sorts every boundary crossing into five risk tiers and sets the authorization each needs, from auto-execution, to an independent machine quorum, to a non-delegable human gate, to hard denial, with a hash-chained audit a third party can check. We calibrate the tiers on 44 machine-readable crossings from the public record, and we evaluate the standard against four authorization structures and four adversaries. Naive voting is not a containment control, while the full protocol executes none of the 32 unsafe crossings, at 22 human escalations, and we say where it still fails. A companion detection harness shows the gated crossings are catchable cheaply, with a rule that fires 66.5 hours before the internal-network pivot, and a companion attestation protocol, ECA, lets a third party verify a run from an exported bundle with no access to the lab's network.
Reviews
The work here is credible, but not excellent. The post-hoc analysis of the single event and demonstration that the protocol works there is certainly some evidence, but seems to overfit to the narrow event in a post-hoc fashion. More consideration of the generalizability would be critical, and the reasoning for many of the asserted claims should be explained more clearly.
It seems quite unclear that the proposed protocol is functional. The paper asserts that "The gated crossings can be caught cheaply and verifiably," but the expense of the detection of the predicates, the costs, and the accuracy are not assessed out of sample, or even in-sample. The list of events found is ad-hoc, and it's unclear that there is a way to automatically detect them other than parsing logs, which is expensive and of unclear accuracy. The list is useful, but far from complete, and partially automatically enforceable, but with an unclear false-positive rate.
The effectiveness is also asserted, not argued for clearly. How does it know whether an identity or login is allowed or not? How does the verifier know if a read even has a side effect? (e.g. HTTP get requests are not actually read-only, and even when thy are, the remote server can take actions that create side effects.) How does the system reliably know if an transaction is high-impact, or prohibited?
The writing was also substandard and messy LLM-generated content; asking a model to reread and rewrite section by section would have improved this, but nothing substitutes for humans carefully reading and ensuring that the work is clear, or better, actually writing or rewriting. (Using the acronym BQP was annoying on top of everything else.)
Read full reviewShow less
I was surprised to see that a naive multi-reviewer vote fails against collusion and outages.
I'd assume it is the fix most teams would reach for first and your paper demonstrates it fails open under both.
I would recommend replacing the scripted reviewers in the adversarial test with the real model from your live loop, since right now your independence claim is only proven against a hand-coded adversary.
This paper addresses a foundational vulnerability in frontier model evaluations and autonomous agent deployments: the complete lack of independent authorization for high-risk boundary crossings during agent runs, as vividly illustrated by the July 2026 intrusion. The author introduces the Boundary Quorum Protocol (BQP) v0.1, a rigorous, risk-proportional authorization framework that gates agent tool execution using tiered risk levels (L0–L4), independence-constrained machine quorums, and non-delegable human gates. Calibrated against 44 documented events from the real incident, BQP is paired with a detection harness demonstrating early warning capabilities and the Evaluation Containment Attestation (ECA) v0.1 bundle for third-party verification. By proving that unsafe crossings can be intercepted with manageable operational overhead and verifiable audit trails, this work delivers a vital blueprint for auditable AI safety governance.
Read full reviewShow less
Cite this project
@misc{tung2026boundary,
title = {{Boundary Quorum Protocol: a verifiable authorization standard for agent evaluations}},
author = {Lam Nguyen Le Tung},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/boundary-quorum-protocol-a-verifiable-authorization-standard-for-agent-evaluations-4hak}},
url = {https://apartresearch.com/sprints/projects/boundary-quorum-protocol-a-verifiable-authorization-standard-for-agent-evaluations-4hak}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …