Participatory Alignment Verification
Shon Pan · Team Shon
Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.
We can improve the fitness for aligned AI versus unaligned AI via costly signaling. An aligned AI can create "signal" by creating tests for misaligned AI, free for itself since it isn't misaligned. A misaligned AI on the other hand, is in a lose-lose situation. Either it needs to create weaker tests for collusion, but this can be detected, or strong tests for misalignment, which would work anyway.
The basic idea is that we give a path for aligned AI to be instrumentally rewarded for alignment, creating a basin for aligned AI.
Reviews
The core idea is clever: force the model to help design the tests that evaluate it, and treat that participation as the signal. I like the principle that participation is allowed, but validation must be external. If done right, this could matter.
Right now, the gap is proof. The demos are interesting, but I want this tested in “hard mode”: an adversarial model explicitly trying to game the protocol, plus real numbers on false positives/negatives. Also, correlated failures across validators feel like the biggest risk and need to be treated as first-class.
Presentation is clear and the idea lands quickly. Execution needs harder stress tests.
this is pretty cool, there's a lot i don't understand about the game theory side of AI control, but the designs seem promising!
If there's actual empirical work done here, i need to see _way_ more details about the scaffold code and experiment methods. But my biggest pet peeve is that i'm genuinely unsure if empirical stuff happened, since case studies sections discuss "proposed plans" as if they're not done yet. At a design level, the reasoning looks sound (at a glance), but my current guess is that there's no empirical work (cuz of the way the case studies section is written-- poorly). This might be fine! theory papers can be great! but I'm annoyed that I'm not sure how to read the case studies section. It's ok to coauthor with claude et al, but be more careful next time.
Cite this project
@misc{pan2026participatory,
title = {{Participatory Alignment Verification}},
author = {Shon Pan},
year = {2026},
month = feb,
note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/participatory-alignment-verification-j9tu}},
url = {https://apartresearch.com/sprints/projects/participatory-alignment-verification-j9tu}
}More from The Technical AI Governance Challenge
- 1st placeView project: LidaSim: Testing AI Policies With Persona-Based Simulations
LidaSim: Testing AI Policies With Persona-Based Simulations
Lida Safety
We simulate well-known figures in AI and politics with agents, scraping large amounts of data to get realistic simulations. Then, we test questions and proposed policies against these public figures, to see which …
- 2nd placeView project: Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
MCL
We present Markov Chain Lock (MCL) watermarking, a cryptographically secure framework for authenticating LLM outputs. MCL constrains token generation to follow a secret Markov chain over SHA-256 vocabulary partitions. …
- 3rd placeView project: Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
AIRAS
The AI safety and governance community is making progress on defining red lines around existential risk from advanced AI systems, and building verification infrastructure to support this objective. However, this is only …