Fault Lines: The Dual-Use AI Governance Vacuum in Asia

Vibha Amarnath

This submission addresses the governance vacuum that emerged in Asian developing states after May 2025, when the US rescinded the AI Diffusion Rule and the Council of Europe's Framework Convention exempted national security AI from treaty scope. The result is a codified asymmetry: developing states have no mechanism to evaluate, contest, or govern the frontier AI operating in their critical infrastructure.

The paper introduces a seven-layer diagnostic framework derived from WMD treaty precedents and applies it across nine countries to show that the vacuum is an architectural failure where every institution with enforcement authority has the wrong jurisdiction for AI, and every institution with AI scope has no enforcement authority. Four confirmed cases document how export controls accelerated unmonitored AI integration into Asian sovereign infrastructure rather than preventing it. Vietnam's Law 134/2025 — more rigorous than any current US federal AI statute sits in the same compute access vacuum as Indonesia, which has enacted nothing.

The submission includes mechanism-level treaty specifications for each of the six absent governance layers, modelled on NPT, CWC, and IAEA precedents, designed as a working resource for the 2027 New York negotiations. The interactive Fault Lines tool (https://fault-lines.netlify.app) makes the full analysis navigable for policy designers and researchers.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is an ambitious and well-presented project addressing a real and important gap: the absence of a multilateral, binding, Global-South-inclusive architecture for dual-use AI governance.

The strongest contribution is the framework it proposes, which decomposes the governance problem into compute/precursor control, capability identification, use-case differentiation, national governance capacity, incident attribution, international verification, and consent/representation.

The live interactive tool is a strong presentation artifact and makes the analysis more usable for policy audiences.

The main weakness is that several load-bearing factual and causal claims seem overstated. The claim that the May 2025 rescission of the AI Diffusion Rule left “no replacement announced” and an ungoverned compute-access vector is too strong. The Malaysia/Huawei case does not truly support the paper’s thesis, at least as written. The paper should try to find a more nuanced description of the governance situation after the rescission. The claims regarding the Council of Europe Convention probably also need correction, at least greater detail/interpretation. Because the “codified asymmetry” claim relies heavily on this treaty, greater precision matters.

Relatedly, references to treaty negotiations should be narrowed and the WMD analogy is under-argued. The paper would be stronger if it specified which WMD governance functions transfer cleanly, which require adaptation, and which may not transfer at all.

The core gap identified by the paper is real and important. But the current version needs to reduce rhetorical certainty and support its causal claims more carefully.

The topic of the project is very timely. Things that could strengthen it include:

(1) Focusing on a narrower subset of the entire problem presented in the submission - currently, both the written document and the built online tool present too much high-level information (e.g. lots of references to policies, historical precedents, etc.) without depth of argumentation. It requires the reader to look for information instead of presenting a scoped out logic explicitly. Essentially, completing an 8-page paper and an online tool during such a short timeframe of a hackathon, as a single person team, seems to have taken away from the submission the ability to offer something deep enough to be concrete. Takeaway: focusing on one of the 4 contribution areas more in detail may have been more beneficial.

(2) The submission would benefit from a deeper exploration of the tensions between dual-use and diffusion, sovereignty, etc. for Asia. Currently, the submission is based on unexplained assumptions about their dynamics.

(3) Unclear audience and this takes away from the impact potential of the submission.

This paper analyzes seven dimensions of past WMD treaties to lay out a potential governance framework for AI. It concludes that current international frameworks have a carve out for AI in national security use cases and developing nations that rely on open-weight models lack governance over AI compute access. The paper also highlights a lack of incident reporting obligations, or audits of cross-border failures, or any mandatory pre-deployment safety evaluations in critical infrastructure. The accompanying website is a useful advocacy tool with great data visualizations and recommendations presented in a compelling manner.

I found the 4 cases meant to support the hypothesis that “Western export restrictions drove systemic integration of unmonitored AI into Asian critical infrastructure” ignore existing domestic and regional governance frameworks in Asia which weakens the argument that there is a “total vacuum” or no operational governance in place. The Malaysia example given seems to contradict the main argument where BIS guidance actually disincentivized the use of Chinese chips. It is not clear from the paper why the law would not apply to “Chinese-origin models” in the case of Singapore or Vietnam, and the argument could be stronger by referring to open-weight models more broadly.

The paper assumes that multi-lateral international treaty-based governance is the best path forward and lays out a potential framework. This works as an advocacy argument for a UN audience, but the limitations section fails to address the reality that negotiating an effective treaty will likely take over a decade, the end result may not benefit the global south, and that investment in domestic or regional solutions to compute governance in the meantime may have more impact. The paper’s claim that effective enforcement mechanisms are crucial can also be strengthened by addressing challenges with treaty ratification and universality given the region’s focus on sovereign AI.

Cite this work

@misc {

title={

(HckPrj) Fault Lines: The Dual-Use AI Governance Vacuum in Asia

},

author={

Vibha Amarnath

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.