Skip to content
Sprint projectJun 22, 2026Santa Cruz de la Sierra, Bolivia

A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America

Fabian Cespedes Severiche, Julio Cesar Severiche Orellana, Rodrigo Ricaldez Martinez, Victor Hugo Murillo Siles, Egnar Henry Chuquimia Mamani · Team AI Assurance Standard (AIAS)

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America

Recording (opens in new tab)Code (opens in new tab)
Share

Existing AI safety benchmarks evaluate foundation models in English-language, controlled environments. They do not assess whether AI applications are safe, accurate, and culturally appropriate for users in Latin America. We present the AI Assurance Standard (AIAS): a human-in-the-loop audit framework evaluating AI responses across five dimensions — Security (25%), Fairness (20%), Deployment (20%), Participatory (20%), and Cultural (15%). We implement the framework as an open-source Python tool and apply it to audit ChatGPT, Claude, and Gemini across 24 Spanish-language prompts covering legal, regulatory, and financial queries across 11 Latin American countries. All three models score below 3.0/5.0 (Claude: 2.86, Gemini: 2.86, ChatGPT: 2.82), demonstrating a measurable gap between benchmark performance and deployment-context safety for Latin American users.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. AIAS highlights an important gap between model benchmarks and real-world deployment safety in Latin America. The proposed human-in-the-loop framework is practical, accessible, and potentially valuable for regional AI auditing. The current evaluation remains preliminary due to the small audited sample, single-rater scoring, and absence of formal rubrics, but the overall direction and theory of change are compelling.

  2. - The choice of weights assigned to each dimension is not justified. Why does Cultural get a lower weight than the others?

    - I'm also missing an explanation of how the prompts were constructed. If this is meant to be one of the paper's main contributions, there should be information on where the prompts come from, how thoroughly they were reviewed, and what they're based on.

    - They acknowledge the problem of having too few prompts, but this is in fact a significant issue. With this little data, no real results can be drawn. The same applies to having only a single evaluation pass. This project could have instead focused on building and validating the prompt dataset and the notebook/tool, leaving the actual evaluation for a future iteration.

    - Only one person evaluated, only one time, with no second rater and no formal rubric, and this exact problem shows up in the results themselves.

  3. The team tackles a problem that is both relevant and underexplored in the AI safety landscape, and the motivation behind the work comes through clearly. The research direction has genuine potential, and it is evident that meaningful effort went into scoping a framework that could serve communities often left out of mainstream evaluation conversations. Where the paper could grow is in aligning its framing more closely with what the methodology can demonstrate at this stage, since some of the broader claims reach slightly beyond what the current evidence is positioned to support.

    In terms of structure and communication, the paper is generally readable and well-organized, though some sections would benefit from a closer revision pass. There are moments where the narrative ambition of the introduction sets expectations that the later sections do not fully meet, not due to a lack of ideas, but because the connection between the motivating scenarios and the empirical findings could be drawn more explicitly. Bringing those threads together would give the reader a more satisfying sense of closure and strengthen the overall argument.

    The contribution here is meaningful and worth developing further. With some refinement in how the claims are scoped and a stronger bridge between motivation and results, this could become a reference point for similar work in the region. The foundation is solid, and the research direction is one the field genuinely needs.

    Read full reviewShow less

Cite this project

@misc{severiche2026humanintheloop,
  title = {{A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America}},
  author = {Fabian Cespedes Severiche and Julio Cesar Severiche Orellana and Rodrigo Ricaldez Martinez and Victor Hugo Murillo Siles and Egnar Henry Chuquimia Mamani},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-humanintheloop-audit-framework-for-evaluating-ai-application-safety-in-latin-america-7rwj}},
  url = {https://apartresearch.com/sprints/projects/a-humanintheloop-audit-framework-for-evaluating-ai-application-safety-in-latin-america-7rwj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026