Latin America Governance & Data Safety Dashboard

Marina Gomes Barbosa, Arthur Lyra Miranda, Thiago Dantas Sousa de Azevedo

Across Latin America, the rapid deployment of AI systems in public services is outpacing the normative frameworks designed to govern them. Existing global tools — such as the OECD.AI Policy Navigator and the IAPP Legislative Tracker — catalog policies at scale but do not produce standardized, verifiable scores that enable direct comparison across countries or distinguish between enacted law and aspirational policy. This gap is particularly consequential for the Global South, where regulatory ambition and enforcement capacity are systematically misaligned.

This paper presents a structured scoring instrument that evaluates eight Latin American countries — Brazil, Argentina, Chile, Colombia, Mexico, Uruguay, Peru, and Bolivia — across three dimensions: Algorithmic Governance (0–4), Data Sovereignty (0–3), and Infrastructure (0–2). Every indicator is grounded in primary official sources and scored through a transparent binary methodology. The tracker is complemented by an AI-powered document analyzer that extracts normative content from uploaded legal texts, classifies it against the indicator framework via a pre-prompted language model, and writes scores directly to the tracker database — reducing what would otherwise require weeks of expert review to a near-instantaneous pipeline.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is a useful and well-motivated contribution. A Latin America-specific, scored, primary-source-traceable AI governance tracker adds some granularity to the regional empirical base beyond existing tools like the OECD.AI Navigator, IAPP and ILIA. Grounding every indicator in a named primary source with a URL, and keeping the scoring binary and contestable, are strengths that make the results independently verifiable. A few changes would meaningfully sharpen it.

First, the categories are not fully distinct, which weakens the composite. Data localization is scored independently of regulated cross-border transfer mechanisms, but these are largely two expressions of the same underlying policy choice about where data may flow, so a country can effectively be credited twice for one decision. The redundancy is compounded by a normative inconsistency: localization counts as a positive point in the score, yet the paper elsewhere treats Bolivia's localization-without-protection as a pathology. Clarifying what each indicator uniquely captures, and whether localization is even a governance positive, would tighten the construct.

Second, the central finding, that governance ambition outpaces enforcement capacity, is a valuable confirmation but not a novel one; it is already well established in the Global South AI governance literature. The paper's real value is the granular, country-level empirical basis it adds beneath that known pattern, and framing the contribution that way (new evidence reinforcing a known dynamic, rather than a new discovery) would set expectations more accurately. Relatedly, the binary scale is in tension with this very finding: it awards the same point to an enacted law, a decree, and a non-binding policy document, so it cannot itself distinguish the intent from the capacity it claims to separate. Colombia scoring 4/4 on a CONPES policy instrument illustrates this. Encoding norm-bindingness directly into the score (the future-work proposal to distinguish law from policy from decree) is what would turn the intent-versus-capacity claim into a measured result.

Finally, the framing should match what was delivered. The abstract and methods foreground an AI-powered document analyzer that reduces weeks of review to an instant pipeline, but the limitations section candidly notes this component was not built or validated and that all scores were produced by manual expert review. The design is ambitious and worth building on for future research, and the honesty is commendable, but the submission has to be judged on the work actually completed, and the delivered results do not match the paper's initial promises. The fix is to present the manual tracker as the contribution (a legitimate and useful one) and the analyzer explicitly as proposed future work. This disclosure should also come much earlier; surfacing it only in the limitations, after four sections have described an automated pipeline, leaves the reader with a picture the paper then has to walk back.

For those of us working on the ground to shape internet policy and digital ethics, this dashboard is a wake-up call. It confirms that high policy scores in our region often reflect mere intent rather than true sovereign capacity. We cannot rely on national AI strategies or ethical frameworks alone; we must demand enforceable laws and localized digital infrastructure. This tool, particularly with its proposed AI-powered document analyzer for rapid legal text classification, will be essential for civil society to hold governments accountable to their stated commitments.

I consider this project a practical and relevant contribution. The dashboard offers a transparent scoring tool for eight Latin American countries across algorithmic governance, data sovereignty, and infrastructure, grounded in primary legal sources. It surfaces an important pattern: governance ambition often appears disconnected from enforcement capacity, as in cases where data localization exists without a broader data protection law. This is a good fit for Global South AI safety.

My main caveats are that the novelty is somewhat overstated and that the AI document analyzer is still a proposed roadmap, not a validated system. The scoring is also binary and somewhat subjective, based on only eight countries, with no check of agreement across reviewers or robustness. To strengthen the project, I would build and validate the analyzer, add evidence on scoring reliability, and sharpen the link to AI safety rather than general governance. Solid and useful policy artifact.

Cite this work

@misc {

title={

(HckPrj) Latin America Governance & Data Safety Dashboard

},

author={

Marina Gomes Barbosa, Arthur Lyra Miranda, Thiago Dantas Sousa de Azevedo

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.