Skip to content
Sprint projectJun 22, 2026Bogotá

HERRAMIENTA DE EVALUACIÓN Y RECOMENDACIÓN PARA LA PROMOCIÓN DE USO RESPONSABLE DE IA EN PYMES LATINOAMERICANAS

Diana Marcela Daza Jaimes, David José Daza Jaimes, Juan Camilo Medina Moreno , Luis Carlos Ordoñez Montenegro , Ángela Pinilla Parra · Team PuenteIA

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: HERRAMIENTA DE EVALUACIÓN Y RECOMENDACIÓN PARA LA PROMOCIÓN DE USO RESPONSABLE DE IA EN PYMES LATINOAMERICANAS

Code (opens in new tab)
Share

Las MiPymes representan el 99,5 % de las unidades productivas de América Latina y el Caribe, pero adoptan la IA generativa de forma acelerada y sin salvaguardas mínimas, en un contexto marcado por la informalidad y la dependencia de proveedores externos. Los estándares globales (NIST AI RMF e ISO/IEC 42001) resultan impracticables a esta escala, y las iniciativas regionales existentes funcionan como listas de verificación estáticas. Este trabajo propone una herramienta tipo SaaS que democratiza la gobernanza ética de la IA al promover la aplicación de normas internacionales y principios fundamentales dentro del sector privado. El instrumento traduce la densidad técnica del NIST AI RMF, la norma ISO/IEC 42001 y un marco ético de cinco principios a un árbol adaptativo de preguntas redactadas en lenguaje llano, y, sobre esa base, despliega tres capas encadenadas (diagnóstico, recomendaciones y ejecución) que culminan en planes de acción para ayudar a nutrir la gobernanza ética de la IA en la empresa.

Al tratarse de un entregable de diseño ilustrado mediante un caso teórico, sus resultados se argumentan en términos de cobertura, robustez y trazabilidad, no de desempeño estadístico. La principal conclusión es que la gobernanza ética de la IA en el Sur Global no debe ser un privilegio corporativo, sino una herramienta de gestión accesible, incluso para las empresas más pequeñas.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is one of the most thoughtfully designed projects in the set. The core architectural decision - route everything with a normative consequence (gap severity, maturity scoring, control mapping) through a fixed deterministic rules engine, and let the LLM enter only at the end to draft the policy - is exactly right, and it directly defeats the "cosmetic compliance" failure mode where two firms with identical answers get different diagnoses on each run. The three-layer anti-hallucination defense (abstention below a 0.35 relevance threshold, automated citation-existence verification, and a second-pass claim-support check) and the hybrid RAG with reciprocal-rank fusion are genuinely well-engineered for a hackathon. Mapping all 30 scorable nodes simultaneously to NIST AI RMF, ISO/IEC 42001, and Floridi's five consolidated principles is solid, and the framing of the problem (99.5% of LATAM productive units are MSMEs for which NIST/ISO are impractical, existing regional efforts are static checklists) is well-evidenced.

    The main limitation is validation, which the team handles with unusual intellectual honesty. (1) There is no field validation - a single constructed theoretical case can demonstrate intended behavior but not external validity. The highest-value next step is a small pilot (even 5-10 real MSMEs) reporting whether the 30/90-day action plans were actually actionable. (2) The 90 node-to-axis weights were LLM-generated and only a sample was hand-audited, yet these weights drive the entire diagnosis; an undetected wrong mapping silently distorts every score. Audit the full set, or at minimum report the audited fraction and the error rate you found in the sample (you note you already caught incorrect control citations - that finding deserves quantification). (3) No inter-rater concordance on the node-to-control/principle mapping; agreement between two domain experts on a subset would substantiate the coverage claim. (4) The RAGAS metrics (recall@6=0.85, citation precision=1.0, faithfulness=0.884) are encouraging but rest on a 10-query golden set - expand it and report per-query so stability is visible.

    Concrete fix: run a small real-MSME pilot and fully audit the frozen weights. The design is strong enough to deserve, and survive, real validation.

    Read full reviewShow less
  2. The project addresses concerns especially salient for Latin America of how small and informal companies can comply with AI risk management frameworks. The focus on these specifically salient risks is smart. The paper is also admirably transparent about its limitations, clear distinguishing what the design demonstrates from what field validation would need to establish. The text sometimes introduces concepts with little contextualization, which undermines its clarity -- such as the reference to the "theory of the Drittwirkung of the Grundrechte" in the Technology, Ethics and Human Rights section. The proposed business model of the SaaS is not discussed, which is a meaningful gap for a tool targeting informal small companies. The walkthrough examples are also hard to follow, though this may partly reflect the machine translation from the original Spanish.

Cite this project

@misc{jaimes2026herramienta,
  title = {{HERRAMIENTA DE EVALUACIÓN Y RECOMENDACIÓN PARA LA PROMOCIÓN DE USO RESPONSABLE DE IA EN PYMES LATINOAMERICANAS}},
  author = {Diana Marcela Daza Jaimes and David José Daza Jaimes and Juan Camilo Medina Moreno and Luis Carlos Ordoñez Montenegro and Ángela Pinilla Parra},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/herramienta-de-evaluacin-y-recomendacin-para-la-promocin-de-uso-responsable-de-ia-en-pymes-latinoamericanas-1gvb}},
  url = {https://apartresearch.com/sprints/projects/herramienta-de-evaluacin-y-recomendacin-para-la-promocin-de-uso-responsable-de-ia-en-pymes-latinoamericanas-1gvb}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026