When Safeguards Stop at the Border, Auditing How OpenAI and Anthropic Allocate Privacy Protections Across Latin American Jurisdictions
Fernanda Stephanie Rokha Sánchez-Umaña, Miguel Andrés Escobar Palta
Analysis of AI privacy policies for LatAm countries based on their personal data protection laws and their comparison with foreign standards using a tool based on RAG architecture and with a judge based on LLM.
This is a strong and useful project. Its main contribution is an auditable method for comparing what AI providers publicly commit to across jurisdictions, rather than simply comparing laws on the books. The distinction between declared protection and actual backend practice is well chosen. A missing policy commitment does not prove that a safeguard is absent in practice, but it does affect what users can see, claim, and contest. The Brazil/Chile contrast is also well motivated, since Brazil has an active data protection authority and an in-force LGPD regime, while Chile’s new law is still in the implementation period.
The empirical finding is narrow but valuable. OpenAI appears to group Brazil and Chile into a general Rest-of-World tier below the EU policy, while Anthropic gives Brazil a dedicated LGPD/ANPD-facing supplement and leaves Chile without an equivalent localized annex. That makes the paper’s central hypothesis plausible: provider localization seems to track in part institutional salience and credible enforcement capacity.
The main limitation is that the validation result is not strong enough to support the most confident versions of the claim. The paper should describe it as something akin to an assistive audit instrument that surfaces candidate discrepancies for human review.
For future work, the most valuable extension would be longitudinal. Chile’s data protection law entering into force creates a natural test: if providers localize their policies after the Chilean agency becomes operational, the institutional-capacity hypothesis becomes much stronger. The project can also add more providers, more jurisdictions, and a technical-behavior layer to compare declared safeguards with actual data controls. Overall, this is a well-scoped and promising governance audit, with a useful method and a defensible core finding, but it should reduce confidence around the classifier and keep its broader claims proportional to the small empirical base.
This is a genuinely good piece of work and the most rigorous of the projects I reviewed. The central finding — that provider safeguards track institutional salience and credible enforcement more than the ambition of the legal text is well-supported by the contrast you build: Brazil earns an ANPD-facing annex from Anthropic while Chile, with a substantively comparable law that simply hasn't come into force yet, stays folded into a generic tier across both providers. The Korean PIPC parallel is a smart addition because it shows the same enforcement-driven localization mechanism operating outside Latin America, which strengthens the abductive reading considerably. And the design choice that carries the whole paper anchoring every verdict to a verbatim snippet a reviewer can locate by text search is exactly what makes the instrument trustworthy and reusable.
I also want to credit the honesty of the limitations section. Naming the self-favouring-judge risk directly (Claude evaluating Anthropic's own policy), acknowledging the retrieval-miss failure mode, and refusing any statistical claim from six probed requirements with only two yielding variation — that restraint is what makes the qualitative finding credible rather than overstated.
Two things would lift it further. First, the dependence on a single judge model remains the load-bearing unresolved risk; even a small re-run of the probe cells with a second, non-Anthropic model would let you show whether the Anthropic-favouring direction is real or absent, and would neutralize the most obvious objection. Second, the engine's 67.9% agreement is doing heavy validation duty, and the gap between that and the 85.7% "right article, right logic" figure deserves more than a sentence — a brief error table showing which of the nine mismatches are doctrinal-threshold disagreements versus genuine misreads would make the "assistive but not deployment-grade" claim concrete. The Chile December 2026 activation as a natural experiment is the right next move and worth foregrounding a longitudinal run is the cleanest available test of your causal claim.
you address a critical and unexamined problem, very good framing, and the instrument you use really is a great contribution! well documented methodology and exceptionally well presented. I also appreciated you addressed the inherent limitations.
Cite this work
@misc {
title={
(HckPrj) When Safeguards Stop at the Border, Auditing How OpenAI and Anthropic Allocate Privacy Protections Across Latin American Jurisdictions
},
author={
Fernanda Stephanie Rokha Sánchez-Umaña, Miguel Andrés Escobar Palta
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


