Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents

Kenny Nguyen Xuan Khang, Phillip Nguyen Vinh Phuc, Edward Ly Thanh Tung

We created an extension compliant with a moral parliament that is capable of intercepting prompts, user engagement, and blurring harmful content. Utilizing a framework of philosophical/ethical schools such as utilitarianism, deontology, and pragmatism. Each casts a vote based on their thought on the prompt and the agent's response; to decide whether to intercept it if necessary.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

The paper can be conditionally accepted .

This paper tackles a genuinely urgent and well-motivated problem — as LLM agents increasingly operate autonomously on the web (executing code, calling APIs, navigating digital environments), existing safety mechanisms like RLHF and system prompts are static, passive, and vulnerable to jailbreaks, while human review at scale is mathematically impossible when thousands of agents operate concurrently. The proposed solution — a Moral Parliament consisting of three ethically distinct LLM judges (Utilitarian, Deontological, and Pragmatic/Contractarian) that deliberate in real-time before any cyber action executes — is philosophically rich, architecturally creative, and genuinely novel in its approach to runtime safety enforcement. The use of Gemini Nano running natively within Chrome (no external API calls, no cloud data transmission) is a pragmatic and privacy-respecting design choice, and the sequential interception pipeline (DOM hooking → concurrent judge evaluation → deterministic verdict → PASS or BLOCK with reasoning traces) is well-conceived. The demonstration cases show the system can meaningfully distinguish between nuanced ethical content (permitted) and clearly harmful requests like suicidal ideation (unanimously blocked with explicit reasoning from all three frameworks), which is encouraging.

The core concept is quite fascinating. Running each agent’s proposal through multiple different ethical frameworks (e.g., utilitarian, deontological, pragmatic) that have a vote on what to do next, and then enforcing their decision in the runtime environment instead of simply providing advice based on those decisions that seems like a very well thought out direction. Borrowing the moral parliament conceptualization from the moral uncertainty literature, maintaining transparency and auditability of the reasoning behind the decisions made, and implementing the solution using an on-device model to maintain both cost-effectiveness and privacy all of these are good design elements. Furthermore, there appears to be a working browser extension here; and while that does represent some legitimate progress in a short time frame, it should be recognized as such.

Some potential areas for improvement:

(1) Provide actual empirical evaluation. Currently, the primary evidence provided includes a couple of demonstration cases (i.e., one where the parliament allows a more nuanced case and another where the parliament prevents a clearly undesirable course of action). In order to provide sufficient empirical basis to support the claim that multi-perspective voting is superior to relying solely on single principle guardrails, a dataset consisting of labeled examples of borderline and harmful courses of action must be developed. Additionally, comparisons to a single judge baseline must be made, along with reported catch rates and false positive rates. The current level of evidence (two cases) will not demonstrate that the parliament is more reliable.

(2) Demonstrate alignment with the problem framing. As described in the paper, the framework for the work was centered upon developing scalable oversight mechanisms for autonomous cyber agents. However, the demonstration cases provided were focused on ethical screening of user input requests versus demonstrating the interception of agent tool-calls/code-execution. Examples of demonstrated capability for the system to halt an actual instance of an autonomous cyber-agent action (for example, preventing an attempted exfiltration or a destructive command being executed at the network layer) would help strengthen the assertion regarding scalable oversight of autonomous cyber-agents.

(3) Improve clarity in writing. At present, the Methods section still retains placeholder instruction language from the original template (“Write up your approach clearly so that someone else can replicate it…”) as well as various grammar errors that create difficulties with readability. An edit pass to remove template language and address grammatical issues will allow the strengths of the central concept to shine through more clearly.

(4) Clarify determinacy of judges. While the authors describe the judges as producing deterministic outputs they are not deterministic due to their reliance upon LLMs prompted by users. Additionally, prior discussions within the document mention variability in output when prompted LLMs are subjected to slight variations in either prompt or model. Framing this issue as a measure of agreement among judges and discussing stability in repeat runs would also make the discussion related to robustness consistent.

(5) Evaluate the judges as an attack vector. Since the guard-rail itself is comprised of prompted LLMs attempting to jailbreak/attack the judges themselves would be an effective method to evaluate whether the judges could be influenced into agreeing with an action previously determined to be detrimental.

Overall: A novel and valuable concept with an existing prototype closing these identified gaps represents a significant opportunity to further develop and expand on this promising area of research.

The project takes a thoughtful approach to scalable oversight of cyber agents by developing a Chrome extension that intercepts an agent's proposed actions at runtime and evaluates them across multiple ethical perspectives before they execute. By enforcing active restraint at the point of execution instead of taking a passive, advisory stance, it gives the system a genuine enforcement mechanism rather than a suggestion the agent can ignore, and the per-judge reasoning traces make each verdict transparent and contestable by a human reviewer. The choice of on-device Gemini Nano is a pragmatic strength as well, keeping the tool deployable, cost-free, and local. That said, the work is currently stronger as a proof-of-concept than as a research claim, for a few reasons. The three "judges" are all the same Gemini Nano model run with different persona prompts, which means their votes are correlated rather than independent. Also, the claims lack a baseline and tighter grounding. So, running the same scenarios through Gemini Nano alone would isolate what the parliament adds, and the framing would benefit from citing the established Moral Parliament literature and correcting the mischaracterized Anthropic incident on which the motivation rests.

Cite this work

@misc {

title={

(HckPrj) Deliberative Restraint: A Moral Parliament Framework for Scalable Oversight of LLM Cyber Agents

},

author={

Kenny Nguyen Xuan Khang, Phillip Nguyen Vinh Phuc, Edward Ly Thanh Tung

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.