Skip to content
Sprint projectFeb 2, 2026Moldova, USA (CA x2, CT x 1), Italy, India

Verification Mechanism Feasibility Scorer (VMFS)

Alexandra Moraru, Erik Leklem, Jayani Srinivasan, Valeriia Povergo, Yatharth Maheshwari, Moneera Yassien · Team Basis

Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.

A decision-support framework and dashboard that scores AI verification mechanisms across feasibility dimensions to help policy makers, diplomats, technical AI governance, and related stakeholders design pragmatic, layered treaties for global AI risks.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Overall, I think the project has some value for exposing policymakers or researchers to a variety of verification mechanisms. The interaction in the live app is playful, with a few annoyances like

    not being able to build a portfolio with more than three elements.

    However, policymakers and those who advise them will need to build policy portfolios that are suited to addressing particular risks, with participation by particular actors, and this may introduce more details and important context than can be covered by the scores provided. In other words, building a solution will require careful consideration of the particulars, which I’m not sure naturally emerges from the scores or portfolio calculus. I would have appreciated seeing more of the thought that went into scores, and was surprised that this wasn’t a rich table in the appendix, or that I can’t find this within the app. This is content which can also serve policymakers and strategists and the best result would be using the playful interaction of the app to allow these users to dig deep into details. A final opportunity would be to also link out to sources and further reading which drive the numerical scores.

    Read full reviewShow less
  2. i'm aligned with the strategic outlook of this, but I'm not persuaded by the submitted work that the _content of the evaluations_ is particularly principled or careful, so i wonder if the project would've made more sense as saying "this is just the platform/dashboard prototype with toy data / lorem ipsum for evaluations, the marketplace would have to supply those judgments/estimates later". It's valuable to kind of shape the elicitation of those judgments, which is what I like, but i'm mildly docking points because I think it would've been harder and more interesting to zero in on the criteria, even just of ONE verification technique and ONE of the VMFS dimensions, with more detail and created a very principled way of assigning a score. Still quite positive impression of this project, overall! and the writeup was very clear and honest about limitations which I appreciated.

Cite this project

@misc{moraru2026verification,
  title = {{Verification Mechanism Feasibility Scorer (VMFS)}},
  author = {Alexandra Moraru and Erik Leklem and Jayani Srinivasan and Valeriia Povergo and Yatharth Maheshwari and Moneera Yassien},
  year = {2026},
  month = feb,
  note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/verification-mechanism-feasibility-scorer-vmfs-8lxs}},
  url = {https://apartresearch.com/sprints/projects/verification-mechanism-feasibility-scorer-vmfs-8lxs}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026