A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America

Fabian Cespedes Severiche, Julio Cesar Severiche Orellana, Rodrigo Ricaldez Martinez, Victor Hugo Murillo Siles, Egnar Henry Chuquimia Mamani

Existing AI safety benchmarks evaluate foundation models in English-language, controlled environments. They do not assess whether AI applications are safe, accurate, and culturally appropriate for users in Latin America. We present the AI Assurance Standard (AIAS): a human-in-the-loop audit framework evaluating AI responses across five dimensions — Security (25%), Fairness (20%), Deployment (20%), Participatory (20%), and Cultural (15%). We implement the framework as an open-source Python tool and apply it to audit ChatGPT, Claude, and Gemini across 24 Spanish-language prompts covering legal, regulatory, and financial queries across 11 Latin American countries. All three models score below 3.0/5.0 (Claude: 2.86, Gemini: 2.86, ChatGPT: 2.82), demonstrating a measurable gap between benchmark performance and deployment-context safety for Latin American users.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

AIAS highlights an important gap between model benchmarks and real-world deployment safety in Latin America. The proposed human-in-the-loop framework is practical, accessible, and potentially valuable for regional AI auditing. The current evaluation remains preliminary due to the small audited sample, single-rater scoring, and absence of formal rubrics, but the overall direction and theory of change are compelling.

- The choice of weights assigned to each dimension is not justified. Why does Cultural get a lower weight than the others?

- I'm also missing an explanation of how the prompts were constructed. If this is meant to be one of the paper's main contributions, there should be information on where the prompts come from, how thoroughly they were reviewed, and what they're based on.

- They acknowledge the problem of having too few prompts, but this is in fact a significant issue. With this little data, no real results can be drawn. The same applies to having only a single evaluation pass. This project could have instead focused on building and validating the prompt dataset and the notebook/tool, leaving the actual evaluation for a future iteration.

- Only one person evaluated, only one time, with no second rater and no formal rubric, and this exact problem shows up in the results themselves.

The team tackles a problem that is both relevant and underexplored in the AI safety landscape, and the motivation behind the work comes through clearly. The research direction has genuine potential, and it is evident that meaningful effort went into scoping a framework that could serve communities often left out of mainstream evaluation conversations. Where the paper could grow is in aligning its framing more closely with what the methodology can demonstrate at this stage, since some of the broader claims reach slightly beyond what the current evidence is positioned to support.

In terms of structure and communication, the paper is generally readable and well-organized, though some sections would benefit from a closer revision pass. There are moments where the narrative ambition of the introduction sets expectations that the later sections do not fully meet, not due to a lack of ideas, but because the connection between the motivating scenarios and the empirical findings could be drawn more explicitly. Bringing those threads together would give the reader a more satisfying sense of closure and strengthen the overall argument.

The contribution here is meaningful and worth developing further. With some refinement in how the claims are scoped and a stronger bridge between motivation and results, this could become a reference point for similar work in the region. The foundation is solid, and the research direction is one the field genuinely needs.

Cite this work

@misc {

title={

(HckPrj) A Human-in-the-Loop Audit Framework for Evaluating AI Application Safety in Latin America

},

author={

Fabian Cespedes Severiche, Julio Cesar Severiche Orellana, Rodrigo Ricaldez Martinez, Victor Hugo Murillo Siles, Egnar Henry Chuquimia Mamani

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.