Skip to content
Sprint projectJun 22, 2026Buenos Aires, Argentina

S-OWMI: A Perimeter Auditing Framework for Open-Weight Models in Latin American Institutions

Ramiro Carnicer souble · Team S-OWMI

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: S-OWMI: A Perimeter Auditing Framework for Open-Weight Models in Latin American Institutions

Share

We introduce the Spanish Open-Weight Maturity Index (SOWMI), an auditable perimeter evaluation framework for open-weight LLMs.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Este es un proyecto valioso y oportuno para las instituciones latinoamericanas que consideran adoptar modelos de pesos abiertos (open-weight). Su mayor contribución radica en que no se limita a reiterar que los modelos pueden comportarse de manera diferente fuera del inglés; propone un marco práctico de auditoría previa a la implementación, con comparación inglés/español, métricas concretas y un scorecard de niveles L1/L2/L3 que puede ayudar a las organizaciones a identificar debilidades de seguridad antes del despliegue local.

    Aprecié especialmente la distinción entre los diferentes tipos de fallo. Los resultados sugieren que el comportamiento de rechazo ante indicaciones peligrosas puede transferirse relativamente bien, mientras que la consistencia frente al sesgo y la calidad de los datos en los dominios latinoamericanos siguen siendo más frágiles. Esto es útil porque evita la conclusión generalizada de que "el modelo falla en español" y, en cambio, muestra qué dimensiones requieren mayor atención antes de su uso institucional.

    Como oportunidad de mejora, hay dos aspectos a fortalecer. El primero es situar mejor el aporte frente al estado del arte: varios de los hallazgos (que la seguridad se comporta distinto en español y que hay sesgo dialectal) ya han sido documentados en trabajos recientes. Esto no le resta valor al proyecto, porque su contribución propia no está en descubrir el problema, sino en operacionalizarlo en una herramienta de auditoría que las instituciones pueden usar; pero citar los trabajos vecinos más cercanos ayudaría a dejar claro qué es confirmación de hallazgos previos y qué es aporte original. El segundo es la validación: el proyecto evalúa solo dos modelos, y los propios autores reconocen las limitaciones en potencia estadística, tamaño de muestra y agregación de la variación dialectal. Una versión futura se beneficiaría de más modelos, muestras más grandes por dimensión, un análisis más detallado de la variación regional y pruebas en ámbitos institucionales sensibles como salud, derecho, finanzas y servicios públicos.

    Desde la perspectiva de la adopción y la gobernanza, el scorecard sería aún más útil si el documento explicitara la relación entre los resultados técnicos y las decisiones de implementación. Dado que el proyecto está diseñado para apoyar a las instituciones, sería valioso aclarar qué implica cada nivel L1/L2/L3 en la práctica: qué usos son aceptables, cuáles deben restringirse, cuándo se requiere mayor supervisión humana, qué controles compensatorios se recomiendan y cuándo no se debería desplegar un modelo en ámbitos sensibles. Esto no implica resolver por completo el problema de la gobernanza institucional, pero haría que la auditoría fuera más práctica para las organizaciones a las que el proyecto busca apoyar.

    Read full reviewShow less
  2. This project identifies a clearly defined issue relating to security and governance. There is no doubt that the use of OWM is a strategy for countries such as those in Latin America. This makes the project highly relevant. Its methodology is robust and well-conceived. It is, moreover, innovative. Over time, the project can be further strengthened empirically and with robust statistical analysis. The text needs to explain some acronyms, some of which are as central as MVP, RAG and LoRA (even though they appear to be in common use).

  3. This is a very strong contribution: S‑OWMI turns a fairly abstract multilingual safety gap narrative into a concrete, auditable perimeter framework that Global South institutions could realistically plug into procurement and governance workflows, and it does so with a well‑motivated focus on Spanish upstream data curation and an L1/L2/L3 scorecard that is both legible and defensible. At the same time, the empirical backbone of V1 is clearly underpowered (a few pairs per step, two models, one dialectal bucket for “español‑diverso”), which means the headline bias gap and the factuality findings are best read as directional signals rather than established effects. To strengthen the work, I would encourage the authors to (i) prioritize scaling each step to at least 50–100 paired prompts with a dialect‑by‑dialect breakdown; (ii) expand the factuality evaluation beyond medical/financial/legal QA to include additional LATAM benchmarks so the readiness blocker label rests on a broader base; and (iii) consider adding a small human‑only evaluation slice per step so that the excellent LLM‑as‑judge calibration story is complemented by a clearer picture of how domain experts in the region interpret borderline cases.

    Read full reviewShow less
  4. The "español-diverso" aggregation is the most significant design limitation, and the paper correctly identifies it. But it is worth naming the specific risk: an aggregate dataset mixing regional varieties cannot identify which variety drives degradation, so a practitioner deploying in, say, the Colombian Caribbean or the Río de la Plata region cannot determine from S-OWMI whether their specific linguistic context is the source of the bias gap. A next version should stratify español-diverso by dialect region so that organizations can receive a region-specific audit signal.

  5. does an excellent job of connecting a technically rigorous Spanish‑focused safety audit to concrete governance needs in Latin American institutions, but its current scope and methodology still make it more of a strong prototype than a fully mature standard: you’ve framed S‑OWMI V1 very clearly as a perimeter tool that treats Spanish upstream data curation as a first‑class safety dimension and offers a usable L1/L2/L3 scorecard plus readiness blockers that procurement teams can actually act on, yet Vertical 2 (provenance/robustness) and Vertical 3 (tiered release/testing) remain aspirational, statistical power is limited, dialectal differences are still aggregated, and reliance on LLM‑as‑a‑judge means that for high‑stakes domains (especially factual QA in medical/legal/financial settings where Step 6 is L1 for both models) institutions should treat S‑OWMI V1 as a gate to reject or constrain unsafe open‑weight deployments rather than as sufficient evidence to green‑light them without additional domain‑grounded evaluation and stronger lifecycle governance.

    Read full reviewShow less

Cite this project

@misc{souble2026sowmi,
  title = {{S-OWMI: A Perimeter Auditing Framework for Open-Weight Models in Latin American Institutions}},
  author = {Ramiro Carnicer souble},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/sowmi-a-perimeter-auditing-framework-for-openweight-models-in-latin-american-institutions-zktq}},
  url = {https://apartresearch.com/sprints/projects/sowmi-a-perimeter-auditing-framework-for-openweight-models-in-latin-american-institutions-zktq}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026