Coverage, Not Faithfulness
Johanna Angulo
CoT monitoring, a leading safety proposal, presupposes that the reasoning it reads is there. Faithfulness metrics do not close that gap: they are defined only on traces that already exist. We measure, from Inspect logs, how often an evaluator can read an agent’s reasoning at all—separating reasoning that was withheld from reasoning that was never produced—and what ceiling that puts on a monitor gated on deliberation. Across 23,541 assistant turns, readability varies more with the serving configuration—the API route, request parameters and one beta header—than across models. In a controlled comparison the same model, task and harness disclose post-tool-result reasoning on 0 of 661 turns versus 667 of 667, depending only on whether that header is sent: a choice the harness makes silently and no evaluation report records. On a complying agent whose reasoning we darken by request flag, a reasoning-only monitor’s recall of harm-pursuing trajectories falls from 0.93 to 0.00 at the first step, while an action-only monitor still catches 0.26 against 0.48 visible. That flag also raised refusals, so the arms are not a clean behavioural A/B. We package this as a reporting standard and an instrument: no rate without the detector’s measured recall, and uninspectable turns counted as failures, not dropped denominators. The contribution is measurement, not a new phenomenon: that CoT can be hidden is known; that a monitor’s measured recall can be an artifact of the evaluation plumbing, moving from 0.93 to 0.00 with no change to the model or the task, is the result.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Coverage, Not Faithfulness
},
author={
Johanna Angulo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


