Skip to content
Sprint projectFeb 1, 2026San Francisco

Participatory Alignment Verification

Shon Pan · Team Shon

Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.

We can improve the fitness for aligned AI versus unaligned AI via costly signaling. An aligned AI can create "signal" by creating tests for misaligned AI, free for itself since it isn't misaligned. A misaligned AI on the other hand, is in a lose-lose situation. Either it needs to create weaker tests for collusion, but this can be detected, or strong tests for misalignment, which would work anyway.

The basic idea is that we give a path for aligned AI to be instrumentally rewarded for alignment, creating a basin for aligned AI.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The core idea is clever: force the model to help design the tests that evaluate it, and treat that participation as the signal. I like the principle that participation is allowed, but validation must be external. If done right, this could matter.

    Right now, the gap is proof. The demos are interesting, but I want this tested in “hard mode”: an adversarial model explicitly trying to game the protocol, plus real numbers on false positives/negatives. Also, correlated failures across validators feel like the biggest risk and need to be treated as first-class.

    Presentation is clear and the idea lands quickly. Execution needs harder stress tests.

  2. this is pretty cool, there's a lot i don't understand about the game theory side of AI control, but the designs seem promising!

    If there's actual empirical work done here, i need to see _way_ more details about the scaffold code and experiment methods. But my biggest pet peeve is that i'm genuinely unsure if empirical stuff happened, since case studies sections discuss "proposed plans" as if they're not done yet. At a design level, the reasoning looks sound (at a glance), but my current guess is that there's no empirical work (cuz of the way the case studies section is written-- poorly). This might be fine! theory papers can be great! but I'm annoyed that I'm not sure how to read the case studies section. It's ok to coauthor with claude et al, but be more careful next time.

Cite this project

@misc{pan2026participatory,
  title = {{Participatory Alignment Verification}},
  author = {Shon Pan},
  year = {2026},
  month = feb,
  note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/participatory-alignment-verification-j9tu}},
  url = {https://apartresearch.com/sprints/projects/participatory-alignment-verification-j9tu}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026