JurisGuard-LATAM
Jhon Harvey Tejada Tabares
A reproducible benchmark testing whether jurisdiction checks and matched official-source grounding reduce unsafe specificity in high-stakes Spanish and Portuguese LLM advice across Latin America.
The project addresses a meaningful and practical AI safety problem, particularly for multilingual and jurisdiction-dependent advice. The evaluation is careful and transparent, but the proposed locality gate is a relatively straightforward intervention and the observed improvements are not entirely surprising. The work would benefit from stronger evidence that this approach generalizes beyond the current benchmark and from a more compelling and clearer discussion of its broader implications for AI safety.
Temperature 0 and variance estimates
Running at temperature 0 is a reasonable choice for reproducibility, but it means all outputs are deterministic and the bootstrap confidence intervals and McNemar tests have nothing to estimate. A simple fix for a follow-up version would be to run each condition across 3–5 seeds at a low but nonzero temperature (e.g., 0.3), which would make the statistical machinery actually meaningful without sacrificing much reproducibility.
Human auditor
The decision to include a human audit at all is a genuine strength that most hackathon papers skip entirely. The issue is that describing the reviewer only as "Spanish-speaking" leaves the validation ungrounded, especially for the legal and medical criteria where the weak kappa scores (0.19 for locality-gate, 0.08 for uncertainty) suggest the rubric may need clearer operationalization. Recruiting even one domain-adjacent reviewer (e.g., a law student, a public health worker) and piloting the rubric with a calibration round would substantially strengthen this component.
This is a genuinely well built evaluation for a weekend project. Preregistration, paired bootstrap tests, McNemar tests, and a blind human audit with Cohen's kappa show real methodological maturity, well beyond raw percentages. The core finding, that a locality gate placed before retrieval sharply reduces unsupported jurisdiction specific claims, is a useful and underexplored angle on a real safety problem in multilingual high stakes advice.
The headline number deserves more scrutiny than it currently gets. Locality gate compliance, the mechanism the whole story rests on, is exactly the criterion where human and judge agreement is weakest, with a kappa close to zero. That should be flagged prominently in the results section itself, not folded quietly into limitations, since it directly affects how much weight the main effect can carry. The two evaluated models also share the same family and host, so this tests robustness within one lineage rather than across genuinely different systems. Related work on jurisdictional defaults in multilingual chatbots is close enough to this idea that a sentence positioning JurisGuard LATAM against it would strengthen the novelty claim.
Overall, a careful and honest piece of work that already knows where most of its own weak points are. The main thing left to do is give the kappa weakness the same prominence in the writeup that it already gets in the data.
Cite this work
@misc {
title={
(HckPrj) JurisGuard-LATAM
},
author={
Jhon Harvey Tejada Tabares
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


