Binding AI Governance in the Global South via Psychometric Metrology
Jesi Martin Maglana · Team-othy
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
As frontier AI systems scale across the Global South, regional regulatory infrastructure has lagged behind. This project proposes a legally defensible AI auditing pipeline for ASEAN state actors by bridging psychometric metrology with regional policy.
Reviews
4/2/4
Criteria 1 - Impact Potential & Innovation: 4
Criteria 2 - Execution Quality: 2
Criteria 3 - Presentation & Clarity: 4
A strong conceptual contribution with no implementation, an unresolved foundational assumption, and results borrowed from its own citations.
The paper raises an important problem -- if AI regulation is going to become binding, regulators need better measurement than simple benchmark percentages, especially in multilingual regions like ASEAN.
Using IRT/CAT as a way to make audits more comparable, cheaper, and less dependent on English-centric benchmarks is a sensible governance direction. The main thing I would improve is concreteness. Since IRT, CAT, and multilingual safety benchmarks already exist, the proposal would be stronger with a small worked example: who maintains the regional item bank, how items are legally validated, how thresholds are set, how different safety dimensions are handled separately, and how model providers can contest audit results. This would make the idea feel less like a high-level metrology proposal and more like an implementable regulatory pathway.
Take the core idea seriously: framing safety evaluation as test-invariant metrology, a score that does not depend on which prompts you happened to sample, is exactly what cross-border enforcement needs, and you tie it well to the live ASEAN window (Vietnam 134/2025, DEFA). The problem is it stays on paper. No calibration, no data, no code, and the eye-catching 99.9% compute number is borrowed from prior work, not shown here. So right now it is a strong proposal, not a result yet tho. The one move that changes that: actually run the SEA-HELM 2PL-IRT calibration you describe, even on a handful of items and two or three models, and report the ability estimates and item parameters. Then deal with multidimensionality head-on (separate scores for biosecurity vs linguistic bias) and test the anomaly filter on real items instead of citing the 84%.
Cite this project
@misc{maglana2026binding,
title = {{Binding AI Governance in the Global South via Psychometric Metrology}},
author = {Jesi Martin Maglana},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/binding-ai-governance-in-the-global-south-via-psychometric-metrology-7xu1}},
url = {https://apartresearch.com/sprints/projects/binding-ai-governance-in-the-global-south-via-psychometric-metrology-7xu1}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …