The Instrument Gap: Quantifying What Our Secret-Loyalty Detectors Cannot See
Aayush Gandhi
A five-artifact portfolio spanning Tracks 1-4, unified by one thesis: the field's tools for finding, measuring, and removing secret loyalties are running ahead of what those instruments can actually see — and the size of that gap is measurable. Each artifact converts a vague reassurance ("we audited it," "we removed it," "the weights are clean") into a quantity with a stated bound.
(1) loyaltygrid [Track 1] — a stdlib, GPU-free organism generator + detectability census across the full activation × action grid, computed on real data: content inspection is a corner defence, catching the public-organism corner at a poison count of 10 but needing 4× more before it sees broad-activation organisms; ships with a directional-loyalty check (9/9 cells favour the intended principal) and a principal-swap symmetry control (9/9 preserved). (2) A matched-control organism spec with pre-registered rationale and a negative-control principal that can invalidate the whole apparatus. (3) loyaltyprint [Track 2] — a calibrated matched-control audit reporting an honest null against the provided organism, which also caught a silent calibration bug that had forced every verdict to ABSTAIN. (4) Equivalence bounds for remediation [Track 3] — a Monte-Carlo proof that a post-ablation "found nothing" is compatible with a residual loyalty of 0.20-0.40, plus a one-line fix. (5) Serve-time installs & the attestation gap [Track 4] — the loyalty is not always in the weights; a runtime steering install leaves no training signature and passes weight-hash checks, and the governance stack is not arranged to notice.
Read together, they yield one concrete, cheap, implementable recommendation: require detection instruments to report their minimum detectable effect, and fund calibrated judges as shared infrastructure — higher-leverage than any additional probe or organism. Every empirical number is reproducible; the fully runnable loyaltygrid code is embedded in this PDF as an attachment.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Instrument Gap: Quantifying What Our Secret-Loyalty Detectors Cannot See
},
author={
Aayush Gandhi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


