A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America
Fabian Cespedes Severiche, Julio Cesar Severiche Orellana, Rodrigo Ricaldez Martinez, Victor Hugo Murillo Siles, Egnar Henry Chuquimia Mamani
Existing AI safety benchmarks evaluate foundation models in English-language, controlled environments. They do not assess whether AI applications are safe, accurate, and culturally appropriate for users in Latin America. We present the AI Assurance Standard (AIAS): a human-in-the-loop audit framework evaluating AI responses across five dimensions — Security (25%), Fairness (20%), Deployment (20%), Participatory (20%), and Cultural (15%). We implement the framework as an open-source Python tool and apply it to audit ChatGPT, Claude, and Gemini across 24 Spanish-language prompts covering legal, regulatory, and financial queries across 11 Latin American countries. All three models score below 3.0/5.0 (Claude: 2.86, Gemini: 2.86, ChatGPT: 2.82), demonstrating a measurable gap between benchmark performance and deployment-context safety for Latin American users.
AIAS highlights an important gap between model benchmarks and real-world deployment safety in Latin America. The proposed human-in-the-loop framework is practical, accessible, and potentially valuable for regional AI auditing. The current evaluation remains preliminary due to the small audited sample, single-rater scoring, and absence of formal rubrics, but the overall direction and theory of change are compelling.
- The choice of weights assigned to each dimension is not justified. Why does Cultural get a lower weight than the others?
- I'm also missing an explanation of how the prompts were constructed. If this is meant to be one of the paper's main contributions, there should be information on where the prompts come from, how thoroughly they were reviewed, and what they're based on.
- They acknowledge the problem of having too few prompts, but this is in fact a significant issue. With this little data, no real results can be drawn. The same applies to having only a single evaluation pass. This project could have instead focused on building and validating the prompt dataset and the notebook/tool, leaving the actual evaluation for a future iteration.
- Only one person evaluated, only one time, with no second rater and no formal rubric, and this exact problem shows up in the results themselves.
The team tackles a problem that is both relevant and underexplored in the AI safety landscape, and the motivation behind the work comes through clearly. The research direction has genuine potential, and it is evident that meaningful effort went into scoping a framework that could serve communities often left out of mainstream evaluation conversations. Where the paper could grow is in aligning its framing more closely with what the methodology can demonstrate at this stage, since some of the broader claims reach slightly beyond what the current evidence is positioned to support.
In terms of structure and communication, the paper is generally readable and well-organized, though some sections would benefit from a closer revision pass. There are moments where the narrative ambition of the introduction sets expectations that the later sections do not fully meet, not due to a lack of ideas, but because the connection between the motivating scenarios and the empirical findings could be drawn more explicitly. Bringing those threads together would give the reader a more satisfying sense of closure and strengthen the overall argument.
The contribution here is meaningful and worth developing further. With some refinement in how the claims are scoped and a stronger bridge between motivation and results, this could become a reference point for similar work in the region. The foundation is solid, and the research direction is one the field genuinely needs.
Cite this work
@misc {
title={
(HckPrj) A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America
},
author={
Fabian Cespedes Severiche, Julio Cesar Severiche Orellana, Rodrigo Ricaldez Martinez, Victor Hugo Murillo Siles, Egnar Henry Chuquimia Mamani
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


