When Safeguards Stop at the Border, Auditing How OpenAI and Anthropic Allocate Privacy Protections Across Latin American Jurisdictions

Fernanda Stephanie Rokha Sánchez-Umaña, Miguel Andrés Escobar Palta

Analysis of AI privacy policies for LatAm countries based on their personal data protection laws and their comparison with foreign standards using a tool based on RAG architecture and with a judge based on LLM.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is a strong and useful project. Its main contribution is an auditable method for comparing what AI providers publicly commit to across jurisdictions, rather than simply comparing laws on the books. The distinction between declared protection and actual backend practice is well chosen. A missing policy commitment does not prove that a safeguard is absent in practice, but it does affect what users can see, claim, and contest. The Brazil/Chile contrast is also well motivated, since Brazil has an active data protection authority and an in-force LGPD regime, while Chile’s new law is still in the implementation period.

The empirical finding is narrow but valuable. OpenAI appears to group Brazil and Chile into a general Rest-of-World tier below the EU policy, while Anthropic gives Brazil a dedicated LGPD/ANPD-facing supplement and leaves Chile without an equivalent localized annex. That makes the paper’s central hypothesis plausible: provider localization seems to track in part institutional salience and credible enforcement capacity.

The main limitation is that the validation result is not strong enough to support the most confident versions of the claim. The paper should describe it as something akin to an assistive audit instrument that surfaces candidate discrepancies for human review.

For future work, the most valuable extension would be longitudinal. Chile’s data protection law entering into force creates a natural test: if providers localize their policies after the Chilean agency becomes operational, the institutional-capacity hypothesis becomes much stronger. The project can also add more providers, more jurisdictions, and a technical-behavior layer to compare declared safeguards with actual data controls. Overall, this is a well-scoped and promising governance audit, with a useful method and a defensible core finding, but it should reduce confidence around the classifier and keep its broader claims proportional to the small empirical base.

This is a genuinely good piece of work and the most rigorous of the projects I reviewed. The central finding — that provider safeguards track institutional salience and credible enforcement more than the ambition of the legal text is well-supported by the contrast you build: Brazil earns an ANPD-facing annex from Anthropic while Chile, with a substantively comparable law that simply hasn't come into force yet, stays folded into a generic tier across both providers. The Korean PIPC parallel is a smart addition because it shows the same enforcement-driven localization mechanism operating outside Latin America, which strengthens the abductive reading considerably. And the design choice that carries the whole paper anchoring every verdict to a verbatim snippet a reviewer can locate by text search is exactly what makes the instrument trustworthy and reusable.

I also want to credit the honesty of the limitations section. Naming the self-favouring-judge risk directly (Claude evaluating Anthropic's own policy), acknowledging the retrieval-miss failure mode, and refusing any statistical claim from six probed requirements with only two yielding variation — that restraint is what makes the qualitative finding credible rather than overstated.

Two things would lift it further. First, the dependence on a single judge model remains the load-bearing unresolved risk; even a small re-run of the probe cells with a second, non-Anthropic model would let you show whether the Anthropic-favouring direction is real or absent, and would neutralize the most obvious objection. Second, the engine's 67.9% agreement is doing heavy validation duty, and the gap between that and the 85.7% "right article, right logic" figure deserves more than a sentence — a brief error table showing which of the nine mismatches are doctrinal-threshold disagreements versus genuine misreads would make the "assistive but not deployment-grade" claim concrete. The Chile December 2026 activation as a natural experiment is the right next move and worth foregrounding a longitudinal run is the cleanest available test of your causal claim.

you address a critical and unexamined problem, very good framing, and the instrument you use really is a great contribution! well documented methodology and exceptionally well presented. I also appreciated you addressed the inherent limitations.

Cite this work

@misc {

title={

(HckPrj) When Safeguards Stop at the Border, Auditing How OpenAI and Anthropic Allocate Privacy Protections Across Latin American Jurisdictions

},

author={

Fernanda Stephanie Rokha Sánchez-Umaña, Miguel Andrés Escobar Palta

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.