Binding AI Governance in the Global South via Psychometric Metrology
Jesi Martin Maglana
As frontier AI systems scale across the Global South, regional regulatory infrastructure has lagged behind. This project proposes a legally defensible AI auditing pipeline for ASEAN state actors by bridging psychometric metrology with regional policy.
4/2/4
Criteria 1 - Impact Potential & Innovation: 4
Criteria 2 - Execution Quality: 2
Criteria 3 - Presentation & Clarity: 4
A strong conceptual contribution with no implementation, an unresolved foundational assumption, and results borrowed from its own citations.
The paper raises an important problem -- if AI regulation is going to become binding, regulators need better measurement than simple benchmark percentages, especially in multilingual regions like ASEAN.
Using IRT/CAT as a way to make audits more comparable, cheaper, and less dependent on English-centric benchmarks is a sensible governance direction. The main thing I would improve is concreteness. Since IRT, CAT, and multilingual safety benchmarks already exist, the proposal would be stronger with a small worked example: who maintains the regional item bank, how items are legally validated, how thresholds are set, how different safety dimensions are handled separately, and how model providers can contest audit results. This would make the idea feel less like a high-level metrology proposal and more like an implementable regulatory pathway.
Take the core idea seriously: framing safety evaluation as test-invariant metrology, a score that does not depend on which prompts you happened to sample, is exactly what cross-border enforcement needs, and you tie it well to the live ASEAN window (Vietnam 134/2025, DEFA). The problem is it stays on paper. No calibration, no data, no code, and the eye-catching 99.9% compute number is borrowed from prior work, not shown here. So right now it is a strong proposal, not a result yet tho. The one move that changes that: actually run the SEA-HELM 2PL-IRT calibration you describe, even on a handful of items and two or three models, and report the ability estimates and item parameters. Then deal with multidimensionality head-on (separate scores for biosecurity vs linguistic bias) and test the anomaly filter on real items instead of citing the 84%.
Cite this work
@misc {
title={
(HckPrj) Binding AI Governance in the Global South via Psychometric Metrology
},
author={
Jesi Martin Maglana
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


