Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents
Kenny Nguyen Xuan Khang, Phillip Nguyen Vinh Phuc, Edward Ly Thanh Tung
We created an extension compliant with a moral parliament that is capable of intercepting prompts, user engagement, and blurring harmful content. Utilizing a framework of philosophical/ethical schools such as utilitarianism, deontology, and pragmatism. Each casts a vote based on their thought on the prompt and the agent's response; to decide whether to intercept it if necessary.
The paper can be conditionally accepted .
This paper tackles a genuinely urgent and well-motivated problem — as LLM agents increasingly operate autonomously on the web (executing code, calling APIs, navigating digital environments), existing safety mechanisms like RLHF and system prompts are static, passive, and vulnerable to jailbreaks, while human review at scale is mathematically impossible when thousands of agents operate concurrently. The proposed solution — a Moral Parliament consisting of three ethically distinct LLM judges (Utilitarian, Deontological, and Pragmatic/Contractarian) that deliberate in real-time before any cyber action executes — is philosophically rich, architecturally creative, and genuinely novel in its approach to runtime safety enforcement. The use of Gemini Nano running natively within Chrome (no external API calls, no cloud data transmission) is a pragmatic and privacy-respecting design choice, and the sequential interception pipeline (DOM hooking → concurrent judge evaluation → deterministic verdict → PASS or BLOCK with reasoning traces) is well-conceived. The demonstration cases show the system can meaningfully distinguish between nuanced ethical content (permitted) and clearly harmful requests like suicidal ideation (unanimously blocked with explicit reasoning from all three frameworks), which is encouraging.
The core concept is quite fascinating. Running each agent’s proposal through multiple different ethical frameworks (e.g., utilitarian, deontological, pragmatic) that have a vote on what to do next, and then enforcing their decision in the runtime environment instead of simply providing advice based on those decisions that seems like a very well thought out direction. Borrowing the moral parliament conceptualization from the moral uncertainty literature, maintaining transparency and auditability of the reasoning behind the decisions made, and implementing the solution using an on-device model to maintain both cost-effectiveness and privacy all of these are good design elements. Furthermore, there appears to be a working browser extension here; and while that does represent some legitimate progress in a short time frame, it should be recognized as such.
Some potential areas for improvement:
(1) Provide actual empirical evaluation. Currently, the primary evidence provided includes a couple of demonstration cases (i.e., one where the parliament allows a more nuanced case and another where the parliament prevents a clearly undesirable course of action). In order to provide sufficient empirical basis to support the claim that multi-perspective voting is superior to relying solely on single principle guardrails, a dataset consisting of labeled examples of borderline and harmful courses of action must be developed. Additionally, comparisons to a single judge baseline must be made, along with reported catch rates and false positive rates. The current level of evidence (two cases) will not demonstrate that the parliament is more reliable.
(2) Demonstrate alignment with the problem framing. As described in the paper, the framework for the work was centered upon developing scalable oversight mechanisms for autonomous cyber agents. However, the demonstration cases provided were focused on ethical screening of user input requests versus demonstrating the interception of agent tool-calls/code-execution. Examples of demonstrated capability for the system to halt an actual instance of an autonomous cyber-agent action (for example, preventing an attempted exfiltration or a destructive command being executed at the network layer) would help strengthen the assertion regarding scalable oversight of autonomous cyber-agents.
(3) Improve clarity in writing. At present, the Methods section still retains placeholder instruction language from the original template (“Write up your approach clearly so that someone else can replicate it…”) as well as various grammar errors that create difficulties with readability. An edit pass to remove template language and address grammatical issues will allow the strengths of the central concept to shine through more clearly.
(4) Clarify determinacy of judges. While the authors describe the judges as producing deterministic outputs they are not deterministic due to their reliance upon LLMs prompted by users. Additionally, prior discussions within the document mention variability in output when prompted LLMs are subjected to slight variations in either prompt or model. Framing this issue as a measure of agreement among judges and discussing stability in repeat runs would also make the discussion related to robustness consistent.
(5) Evaluate the judges as an attack vector. Since the guard-rail itself is comprised of prompted LLMs attempting to jailbreak/attack the judges themselves would be an effective method to evaluate whether the judges could be influenced into agreeing with an action previously determined to be detrimental.
Overall: A novel and valuable concept with an existing prototype closing these identified gaps represents a significant opportunity to further develop and expand on this promising area of research.
The project takes a thoughtful approach to scalable oversight of cyber agents by developing a Chrome extension that intercepts an agent's proposed actions at runtime and evaluates them across multiple ethical perspectives before they execute. By enforcing active restraint at the point of execution instead of taking a passive, advisory stance, it gives the system a genuine enforcement mechanism rather than a suggestion the agent can ignore, and the per-judge reasoning traces make each verdict transparent and contestable by a human reviewer. The choice of on-device Gemini Nano is a pragmatic strength as well, keeping the tool deployable, cost-free, and local. That said, the work is currently stronger as a proof-of-concept than as a research claim, for a few reasons. The three "judges" are all the same Gemini Nano model run with different persona prompts, which means their votes are correlated rather than independent. Also, the claims lack a baseline and tighter grounding. So, running the same scenarios through Gemini Nano alone would isolate what the parliament adds, and the framing would benefit from citing the established Moral Parliament literature and correcting the mischaracterized Anthropic incident on which the motivation rests.
Cite this work
@misc {
title={
(HckPrj) Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents
},
author={
Kenny Nguyen Xuan Khang, Phillip Nguyen Vinh Phuc, Edward Ly Thanh Tung
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


