Coverage, Not Faithfulness
Johanna Angulo · Team JAJA
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
CoT monitoring, a leading safety proposal, presupposes that the reasoning it reads is there. Faithfulness metrics do not close that gap: they are defined only on traces that already exist. We measure, from Inspect logs, how often an evaluator can read an agent’s reasoning at all—separating reasoning that was withheld from reasoning that was never produced—and what ceiling that puts on a monitor gated on deliberation. Across 23,541 assistant turns, readability varies more with the serving configuration—the API route, request parameters and one beta header—than across models. In a controlled comparison the same model, task and harness disclose post-tool-result reasoning on 0 of 661 turns versus 667 of 667, depending only on whether that header is sent: a choice the harness makes silently and no evaluation report records. On a complying agent whose reasoning we darken by request flag, a reasoning-only monitor’s recall of harm-pursuing trajectories falls from 0.93 to 0.00 at the first step, while an action-only monitor still catches 0.26 against 0.48 visible. That flag also raised refusals, so the arms are not a clean behavioural A/B. We package this as a reporting standard and an instrument: no rate without the detector’s measured recall, and uninspectable turns counted as failures, not dropped denominators. The contribution is measurement, not a new phenomenon: that CoT can be hidden is known; that a monitor’s measured recall can be an artifact of the evaluation plumbing, moving from 0.93 to 0.00 with no change to the model or the task, is the result.

Reviews
This project addresses a useful question: whether a reasoning-based safety monitor can access the information it needs.
The distinction between produced and readable reasoning, explicit denominators, and released code are valuable contributions. The main improvement is to align the reported methods with the implementation, particularly provider routing and the definition of positive trajectories.
Comparing identical trajectories with reasoning visible versus masked, would better isolate visibility effects. A simpler abstract and the referenced figures would also make the findings easier to follow.
I think this project highlights a useful concern for teams that monitor an agent’s written reasoning. The paper compares settings for the same model and reports a substantial difference in how much reasoning was returned after tool use. Logging those settings alongside evaluation results seems like a practical contribution.
I am less convinced by the claim that the amount of visible reasoning sets a general limit on detection. The overall amount of visible reasoning does not tell us whether it was available in the cases where the agent became misaligned.
The detection analysis also changes which runs are included as they finish, making the results harder to interpret. I would keep the finding that request settings change how much reasoning is visible as the main contribution, and narrow the broader claim on detection. Testing a monitor on the same recorded runs, with different amounts of reasoning hidden, would give a clearer picture of what it misses.
Read full reviewShow less
Cite this project
@misc{angulo2026coverage,
title = {{Coverage, Not Faithfulness}},
author = {Johanna Angulo},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/coverage-not-faithfulness-rna6}},
url = {https://apartresearch.com/sprints/projects/coverage-not-faithfulness-rna6}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …