Verification Mechanism Feasibility Scorer (VMFS)
Alexandra Moraru, Erik Leklem, Jayani Srinivasan, Valeriia Povergo, Yatharth Maheshwari, Moneera Yassien · Team Basis
Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.
A decision-support framework and dashboard that scores AI verification mechanisms across feasibility dimensions to help policy makers, diplomats, technical AI governance, and related stakeholders design pragmatic, layered treaties for global AI risks.
Reviews
Overall, I think the project has some value for exposing policymakers or researchers to a variety of verification mechanisms. The interaction in the live app is playful, with a few annoyances like
not being able to build a portfolio with more than three elements.
However, policymakers and those who advise them will need to build policy portfolios that are suited to addressing particular risks, with participation by particular actors, and this may introduce more details and important context than can be covered by the scores provided. In other words, building a solution will require careful consideration of the particulars, which I’m not sure naturally emerges from the scores or portfolio calculus. I would have appreciated seeing more of the thought that went into scores, and was surprised that this wasn’t a rich table in the appendix, or that I can’t find this within the app. This is content which can also serve policymakers and strategists and the best result would be using the playful interaction of the app to allow these users to dig deep into details. A final opportunity would be to also link out to sources and further reading which drive the numerical scores.
Read full reviewShow less
i'm aligned with the strategic outlook of this, but I'm not persuaded by the submitted work that the _content of the evaluations_ is particularly principled or careful, so i wonder if the project would've made more sense as saying "this is just the platform/dashboard prototype with toy data / lorem ipsum for evaluations, the marketplace would have to supply those judgments/estimates later". It's valuable to kind of shape the elicitation of those judgments, which is what I like, but i'm mildly docking points because I think it would've been harder and more interesting to zero in on the criteria, even just of ONE verification technique and ONE of the VMFS dimensions, with more detail and created a very principled way of assigning a score. Still quite positive impression of this project, overall! and the writeup was very clear and honest about limitations which I appreciated.
Cite this project
@misc{moraru2026verification,
title = {{Verification Mechanism Feasibility Scorer (VMFS)}},
author = {Alexandra Moraru and Erik Leklem and Jayani Srinivasan and Valeriia Povergo and Yatharth Maheshwari and Moneera Yassien},
year = {2026},
month = feb,
note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/verification-mechanism-feasibility-scorer-vmfs-8lxs}},
url = {https://apartresearch.com/sprints/projects/verification-mechanism-feasibility-scorer-vmfs-8lxs}
}More from The Technical AI Governance Challenge
- 1st placeView project: LidaSim: Testing AI Policies With Persona-Based Simulations
LidaSim: Testing AI Policies With Persona-Based Simulations
Lida Safety
We simulate well-known figures in AI and politics with agents, scraping large amounts of data to get realistic simulations. Then, we test questions and proposed policies against these public figures, to see which …
- 2nd placeView project: Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
MCL
We present Markov Chain Lock (MCL) watermarking, a cryptographically secure framework for authenticating LLM outputs. MCL constrains token generation to follow a secret Markov chain over SHA-256 vocabulary partitions. …
- 3rd placeView project: Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
AIRAS
The AI safety and governance community is making progress on defining red lines around existential risk from advanced AI systems, and building verification infrastructure to support this objective. However, this is only …