Skip to content
Sprint projectSep 13, 2026Valencia

Coverage, Not Faithfulness

Johanna Angulo · Team JAJA

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

CoT monitoring, a leading safety proposal, presupposes that the reasoning it reads is there. Faithfulness metrics do not close that gap: they are defined only on traces that already exist. We measure, from Inspect logs, how often an evaluator can read an agent’s reasoning at all—separating reasoning that was withheld from reasoning that was never produced—and what ceiling that puts on a monitor gated on deliberation. Across 23,541 assistant turns, readability varies more with the serving configuration—the API route, request parameters and one beta header—than across models. In a controlled comparison the same model, task and harness disclose post-tool-result reasoning on 0 of 661 turns versus 667 of 667, depending only on whether that header is sent: a choice the harness makes silently and no evaluation report records. On a complying agent whose reasoning we darken by request flag, a reasoning-only monitor’s recall of harm-pursuing trajectories falls from 0.93 to 0.00 at the first step, while an action-only monitor still catches 0.26 against 0.48 visible. That flag also raised refusals, so the arms are not a clean behavioural A/B. We package this as a reporting standard and an instrument: no rate without the detector’s measured recall, and uninspectable turns counted as failures, not dropped denominators. The contribution is measurement, not a new phenomenon: that CoT can be hidden is known; that a monitor’s measured recall can be an artifact of the evaluation plumbing, moving from 0.93 to 0.00 with no change to the model or the task, is the result.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project addresses a useful question: whether a reasoning-based safety monitor can access the information it needs.

    The distinction between produced and readable reasoning, explicit denominators, and released code are valuable contributions. The main improvement is to align the reported methods with the implementation, particularly provider routing and the definition of positive trajectories.

    Comparing identical trajectories with reasoning visible versus masked, would better isolate visibility effects. A simpler abstract and the referenced figures would also make the findings easier to follow.

  2. I think this project highlights a useful concern for teams that monitor an agent’s written reasoning. The paper compares settings for the same model and reports a substantial difference in how much reasoning was returned after tool use. Logging those settings alongside evaluation results seems like a practical contribution.

    I am less convinced by the claim that the amount of visible reasoning sets a general limit on detection. The overall amount of visible reasoning does not tell us whether it was available in the cases where the agent became misaligned.

    The detection analysis also changes which runs are included as they finish, making the results harder to interpret. I would keep the finding that request settings change how much reasoning is visible as the main contribution, and narrow the broader claim on detection. Testing a monitor on the same recorded runs, with different amounts of reasoning hidden, would give a clearer picture of what it misses.

    Read full reviewShow less

Cite this project

@misc{angulo2026coverage,
  title = {{Coverage, Not Faithfulness}},
  author = {Johanna Angulo},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/coverage-not-faithfulness-rna6}},
  url = {https://apartresearch.com/sprints/projects/coverage-not-faithfulness-rna6}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026