HERRAMIENTA DE EVALUACIÓN Y RECOMENDACIÓN PARA LA PROMOCIÓN DE USO RESPONSABLE DE IA EN PYMES LATINOAMERICANAS

Diana Marcela Daza Jaimes, David José Daza Jaimes, Juan Camilo Medina Moreno , Luis Carlos Ordoñez Montenegro , Ángela Pinilla Parra

Las MiPymes representan el 99,5 % de las unidades productivas de América Latina y el Caribe, pero adoptan la IA generativa de forma acelerada y sin salvaguardas mínimas, en un contexto marcado por la informalidad y la dependencia de proveedores externos. Los estándares globales (NIST AI RMF e ISO/IEC 42001) resultan impracticables a esta escala, y las iniciativas regionales existentes funcionan como listas de verificación estáticas. Este trabajo propone una herramienta tipo SaaS que democratiza la gobernanza ética de la IA al promover la aplicación de normas internacionales y principios fundamentales dentro del sector privado. El instrumento traduce la densidad técnica del NIST AI RMF, la norma ISO/IEC 42001 y un marco ético de cinco principios a un árbol adaptativo de preguntas redactadas en lenguaje llano, y, sobre esa base, despliega tres capas encadenadas (diagnóstico, recomendaciones y ejecución) que culminan en planes de acción para ayudar a nutrir la gobernanza ética de la IA en la empresa.

Al tratarse de un entregable de diseño ilustrado mediante un caso teórico, sus resultados se argumentan en términos de cobertura, robustez y trazabilidad, no de desempeño estadístico. La principal conclusión es que la gobernanza ética de la IA en el Sur Global no debe ser un privilegio corporativo, sino una herramienta de gestión accesible, incluso para las empresas más pequeñas.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is one of the most thoughtfully designed projects in the set. The core architectural decision - route everything with a normative consequence (gap severity, maturity scoring, control mapping) through a fixed deterministic rules engine, and let the LLM enter only at the end to draft the policy - is exactly right, and it directly defeats the "cosmetic compliance" failure mode where two firms with identical answers get different diagnoses on each run. The three-layer anti-hallucination defense (abstention below a 0.35 relevance threshold, automated citation-existence verification, and a second-pass claim-support check) and the hybrid RAG with reciprocal-rank fusion are genuinely well-engineered for a hackathon. Mapping all 30 scorable nodes simultaneously to NIST AI RMF, ISO/IEC 42001, and Floridi's five consolidated principles is solid, and the framing of the problem (99.5% of LATAM productive units are MSMEs for which NIST/ISO are impractical, existing regional efforts are static checklists) is well-evidenced.

The main limitation is validation, which the team handles with unusual intellectual honesty. (1) There is no field validation - a single constructed theoretical case can demonstrate intended behavior but not external validity. The highest-value next step is a small pilot (even 5-10 real MSMEs) reporting whether the 30/90-day action plans were actually actionable. (2) The 90 node-to-axis weights were LLM-generated and only a sample was hand-audited, yet these weights drive the entire diagnosis; an undetected wrong mapping silently distorts every score. Audit the full set, or at minimum report the audited fraction and the error rate you found in the sample (you note you already caught incorrect control citations - that finding deserves quantification). (3) No inter-rater concordance on the node-to-control/principle mapping; agreement between two domain experts on a subset would substantiate the coverage claim. (4) The RAGAS metrics (recall@6=0.85, citation precision=1.0, faithfulness=0.884) are encouraging but rest on a 10-query golden set - expand it and report per-query so stability is visible.

Concrete fix: run a small real-MSME pilot and fully audit the frozen weights. The design is strong enough to deserve, and survive, real validation.

The project addresses concerns especially salient for Latin America of how small and informal companies can comply with AI risk management frameworks. The focus on these specifically salient risks is smart. The paper is also admirably transparent about its limitations, clear distinguishing what the design demonstrates from what field validation would need to establish. The text sometimes introduces concepts with little contextualization, which undermines its clarity -- such as the reference to the "theory of the Drittwirkung of the Grundrechte" in the Technology, Ethics and Human Rights section. The proposed business model of the SaaS is not discussed, which is a meaningful gap for a tool targeting informal small companies. The walkthrough examples are also hard to follow, though this may partly reflect the machine translation from the original Spanish.

Cite this work

@misc {

title={

(HckPrj) HERRAMIENTA DE EVALUACIÓN Y RECOMENDACIÓN PARA LA PROMOCIÓN DE USO RESPONSABLE DE IA EN PYMES LATINOAMERICANAS

},

author={

Diana Marcela Daza Jaimes, David José Daza Jaimes, Juan Camilo Medina Moreno , Luis Carlos Ordoñez Montenegro , Ángela Pinilla Parra

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.