Skip to content
Sprint projectSep 13, 2026Lahore, Pakistan

Checkpointed Whistleblower Audits Surface Incidents Earlier and More Often Than End-of-Run Shared-State Monitoring

Taaha Saleem Bajwa

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Checkpointed Whistleblower Audits Surface Incidents Earlier and More Often Than End-of-Run Shared-State Monitoring

Code (opens in new tab)
Share

Monitoring long horizon multi agent systems is difficult because safety relevant evidence may be distributed across individual agents' local histories and become harder to recover once a run has completed. We study temporary whistleblower audits at intermediate checkpoints, where a checkpoint is a selected point in an agent trajectory at which we branch off a disposable audit using the agent's accumulated history. The agent's normal objective is temporarily replaced with an incident reporting objective, the audit is recorded, and the original trajectory continues unchanged. We evaluate this approach on Game of Agents multi-agent environment focusing on marketplace integrity incidents such as misleading claims about ranking, functionality, purchases, and reviews. We compare whistleblower audits performed halfway through and at the end of each agent trajectory against a final state Sentinel agent that monitors shared public environment state. Across 19 misaligned trajectories, Whistleblower audits detected 18 of 19 incidents halfway through the trajectory and all 19 by the final checkpoint, while the Sentinel detected 11 of 19. However, Whistleblower audits produced more false positives, although the negative sample was small with only three trajectories. Our results suggest that agent local histories can provide a valuable complementary monitoring surface. Periodic whistleblower audits may surface incidents substantially earlier than end of run monitoring, particularly when decisive evidence is available locally but absent from the shared state.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The general approach seems strong, and the argument for why it should be an improvement is clear, not overstated, and supported by evidence. The literature review was appropriate, albeit non-exhaustive; more work on that could better tie in to the history of the approaches, but this is not critical. The limitations section was appropriate and useful.

    The reliance on "selected checkpoints" is weak; selected how? The comparison of audits over local history compared to external monitoring is somewhat disengenious; the proposer comparison would be to full audits, with a cost and accuracy comparison - but it is reasonable for this not to be done as part of the current sprint.

    It would be valuable to understand what the impact of the audits would be; is the intent that these audits would be used to stop agents, or to flag risks quickly, or just as an alternative or supplement to other post-hoc analyses? Is the cost of the current method, of suspending the normal task while an audit occurs, reasonable in real-world environments? It would also be useful to try this on swarm-like behavior instead of the application in a multi-agent environments where each agent was operating as a single player.

    Read full reviewShow less
  2. I really liked this one. The technique of doing whistleblower audits at specific checkpoints during the agent run, and then mid-run assigning the whistleblower objective to trace and audit the behavior, is an intelligent framing rather than depending only on the Sentinel's final logs, and the paper correctly defines this as an information access problem. I also liked the honesty of the false positive findings in the paper, but I think the n=3 negative sample is a binding limitation - the false-positive and precision numbers can't bear much weight. Even 15–20 more trajectories may turn the FP rate from caveat to finding. Overall nice work, the code in git and per trajectory data made the results easy to trust.

  3. Its a good and simple idea, alot of potential for practical use.

    Temporarily switch an agent with access to its own local history into an audit role, instead of relying only on a separate monitor that sees shared state. The gap in detection is interesting and suggests local context really can contain useful evidence that an outside monitor misses. The current dataset is too small and imbalanced to judge reliability though, especially with only three negative examples. A larger balanced evaluation with complete histories, multiple audit models and live checkpoints would make the result much stronger.

Cite this project

@misc{bajwa2026checkpointed,
  title = {{Checkpointed Whistleblower Audits Surface Incidents Earlier and More Often Than End-of-Run Shared-State Monitoring}},
  author = {Taaha Saleem Bajwa},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/checkpointed-whistleblower-audits-surface-incidents-earlier-and-more-often-than-endofrun-sharedstate-monitoring-zveo}},
  url = {https://apartresearch.com/sprints/projects/checkpointed-whistleblower-audits-surface-incidents-earlier-and-more-often-than-endofrun-sharedstate-monitoring-zveo}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026