Skip to content
Sprint projectJun 21, 2026Ho Chi Minh

Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents

Kenny Nguyen Xuan Khang, Phillip Nguyen Vinh Phuc, Edward Ly Thanh Tung · Team PhilanthropistEnjoyers

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents

Presentation

Presentation: Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents

Code (opens in new tab)
Share

We created an extension compliant with a moral parliament that is capable of intercepting prompts, user engagement, and blurring harmful content. Utilizing a framework of philosophical/ethical schools such as utilitarianism, deontology, and pragmatism. Each casts a vote based on their thought on the prompt and the agent's response; to decide whether to intercept it if necessary.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The paper can be conditionally accepted .

    This paper tackles a genuinely urgent and well-motivated problem — as LLM agents increasingly operate autonomously on the web (executing code, calling APIs, navigating digital environments), existing safety mechanisms like RLHF and system prompts are static, passive, and vulnerable to jailbreaks, while human review at scale is mathematically impossible when thousands of agents operate concurrently. The proposed solution — a Moral Parliament consisting of three ethically distinct LLM judges (Utilitarian, Deontological, and Pragmatic/Contractarian) that deliberate in real-time before any cyber action executes — is philosophically rich, architecturally creative, and genuinely novel in its approach to runtime safety enforcement. The use of Gemini Nano running natively within Chrome (no external API calls, no cloud data transmission) is a pragmatic and privacy-respecting design choice, and the sequential interception pipeline (DOM hooking → concurrent judge evaluation → deterministic verdict → PASS or BLOCK with reasoning traces) is well-conceived. The demonstration cases show the system can meaningfully distinguish between nuanced ethical content (permitted) and clearly harmful requests like suicidal ideation (unanimously blocked with explicit reasoning from all three frameworks), which is encouraging.

    Read full reviewShow less
  2. The core concept is quite fascinating. Running each agent’s proposal through multiple different ethical frameworks (e.g., utilitarian, deontological, pragmatic) that have a vote on what to do next, and then enforcing their decision in the runtime environment instead of simply providing advice based on those decisions that seems like a very well thought out direction. Borrowing the moral parliament conceptualization from the moral uncertainty literature, maintaining transparency and auditability of the reasoning behind the decisions made, and implementing the solution using an on-device model to maintain both cost-effectiveness and privacy all of these are good design elements. Furthermore, there appears to be a working browser extension here; and while that does represent some legitimate progress in a short time frame, it should be recognized as such.

    Some potential areas for improvement:

    (1) Provide actual empirical evaluation. Currently, the primary evidence provided includes a couple of demonstration cases (i.e., one where the parliament allows a more nuanced case and another where the parliament prevents a clearly undesirable course of action). In order to provide sufficient empirical basis to support the claim that multi-perspective voting is superior to relying solely on single principle guardrails, a dataset consisting of labeled examples of borderline and harmful courses of action must be developed. Additionally, comparisons to a single judge baseline must be made, along with reported catch rates and false positive rates. The current level of evidence (two cases) will not demonstrate that the parliament is more reliable.

    (2) Demonstrate alignment with the problem framing. As described in the paper, the framework for the work was centered upon developing scalable oversight mechanisms for autonomous cyber agents. However, the demonstration cases provided were focused on ethical screening of user input requests versus demonstrating the interception of agent tool-calls/code-execution. Examples of demonstrated capability for the system to halt an actual instance of an autonomous cyber-agent action (for example, preventing an attempted exfiltration or a destructive command being executed at the network layer) would help strengthen the assertion regarding scalable oversight of autonomous cyber-agents.

    (3) Improve clarity in writing. At present, the Methods section still retains placeholder instruction language from the original template (“Write up your approach clearly so that someone else can replicate it…”) as well as various grammar errors that create difficulties with readability. An edit pass to remove template language and address grammatical issues will allow the strengths of the central concept to shine through more clearly.

    (4) Clarify determinacy of judges. While the authors describe the judges as producing deterministic outputs they are not deterministic due to their reliance upon LLMs prompted by users. Additionally, prior discussions within the document mention variability in output when prompted LLMs are subjected to slight variations in either prompt or model. Framing this issue as a measure of agreement among judges and discussing stability in repeat runs would also make the discussion related to robustness consistent.

    (5) Evaluate the judges as an attack vector. Since the guard-rail itself is comprised of prompted LLMs attempting to jailbreak/attack the judges themselves would be an effective method to evaluate whether the judges could be influenced into agreeing with an action previously determined to be detrimental.

    Overall: A novel and valuable concept with an existing prototype closing these identified gaps represents a significant opportunity to further develop and expand on this promising area of research.

    Read full reviewShow less
  3. The project takes a thoughtful approach to scalable oversight of cyber agents by developing a Chrome extension that intercepts an agent's proposed actions at runtime and evaluates them across multiple ethical perspectives before they execute. By enforcing active restraint at the point of execution instead of taking a passive, advisory stance, it gives the system a genuine enforcement mechanism rather than a suggestion the agent can ignore, and the per-judge reasoning traces make each verdict transparent and contestable by a human reviewer. The choice of on-device Gemini Nano is a pragmatic strength as well, keeping the tool deployable, cost-free, and local. That said, the work is currently stronger as a proof-of-concept than as a research claim, for a few reasons. The three "judges" are all the same Gemini Nano model run with different persona prompts, which means their votes are correlated rather than independent. Also, the claims lack a baseline and tighter grounding. So, running the same scenarios through Gemini Nano alone would isolate what the parliament adds, and the framing would benefit from citing the established Moral Parliament literature and correcting the mischaracterized Anthropic incident on which the motivation rests.

    Read full reviewShow less

Cite this project

@misc{khang2026deliberative,
  title = {{Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents}},
  author = {Kenny Nguyen Xuan Khang and Phillip Nguyen Vinh Phuc and Edward Ly Thanh Tung},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/deliberative-restraint-a-moral-parliament-framework-for-scalable-oversight-of-llm-cyber-agents-iq0a}},
  url = {https://apartresearch.com/sprints/projects/deliberative-restraint-a-moral-parliament-framework-for-scalable-oversight-of-llm-cyber-agents-iq0a}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026