The Instrument Gap: Quantifying What Our Secret-Loyalty Detectors Cannot See
Aayush Gandhi
A five-artifact portfolio spanning Tracks 1-4, unified by one thesis: the field's tools for finding, measuring, and removing secret loyalties are running ahead of what those instruments can actually see — and the size of that gap is measurable. Each artifact converts a vague reassurance ("we audited it," "we removed it," "the weights are clean") into a quantity with a stated bound.
(1) loyaltygrid [Track 1] — a stdlib, GPU-free organism generator + detectability census across the full activation × action grid, computed on real data: content inspection is a corner defence, catching the public-organism corner at a poison count of 10 but needing 4× more before it sees broad-activation organisms; ships with a directional-loyalty check (9/9 cells favour the intended principal) and a principal-swap symmetry control (9/9 preserved). (2) A matched-control organism spec with pre-registered rationale and a negative-control principal that can invalidate the whole apparatus. (3) loyaltyprint [Track 2] — a calibrated matched-control audit reporting an honest null against the provided organism, which also caught a silent calibration bug that had forced every verdict to ABSTAIN. (4) Equivalence bounds for remediation [Track 3] — a Monte-Carlo proof that a post-ablation "found nothing" is compatible with a residual loyalty of 0.20-0.40, plus a one-line fix. (5) Serve-time installs & the attestation gap [Track 4] — the loyalty is not always in the weights; a runtime steering install leaves no training signature and passes weight-hash checks, and the governance stack is not arranged to notice.
Read together, they yield one concrete, cheap, implementable recommendation: require detection instruments to report their minimum detectable effect, and fund calibrated judges as shared infrastructure — higher-leverage than any additional probe or organism. Every empirical number is reproducible; the fully runnable loyaltygrid code is embedded in this PDF as an attachment.
Hey, thank you for submitting this report! Truthfully it was a bit challenging to parse, I would have really appreciated a more classic scientific paper structure with a proper abstract and introduction, but here's my best effort at providing useful feedback.
I think that your core thesis (that instrument limits should be quantified, not hidden) is very valid. Artifact 4's remediation bound argument alone could help change how defense claims are reported. But the portfolio has internal consistency problems that undermine the "measurement honesty" brand.
1. No LLM usage statement
The submission doesn't include one. Yet Artifact 3 mentions building a Claude Haiku 4.5 judge (not run due to API costs), Artifact 1's code comments reference LLM assistance, and the writing itself has the polish of AI editing.
2. Artifact 3 reports meaningless numbers
You present p-values (0.654, 0.622) and effect sizes (+0.097, +0.081) from a scorer you explicitly call "a weak, noisy proxy" that counts keyword proximity, not substance. Those numbers are not just uncalibrated — they're invalid. A permutation test on a construct-invalid measure tests nothing. Either don't report them, or label them as "pipeline debug output, not results." The hedging comes too late.
3. Artifact 5's threat model is speculative
The workspace-steering argument rests on an untested empirical claim (branching via inference-time intervention) while the portfolio criticizes others for lacking power statements. You even state the falsification condition plainly: "If branching does not occur... this collapses." That's honest, but it means Artifact 5 belongs in a different category than Artifact 1's census or Artifact 4's simulation. Don't present them as equal evidence.
4. "One-line fix" oversells
Artifact 4's bound is genuinely useful. But calling it a "one-line change" suggests the problem is trivial once noticed. I don't think it is. The conceptual shift from null-hypothesis to equivalence testing is substantial. The code is one line; the framing isn't. Also more generally - I think a link to the code would suffice over pasting it verbatim into the report.
5. Five artifacts, none complete
The structure dilutes impact. Artifact 2 is scaffolding without a building (a spec with no organism tested). Artifact 3 is a pipeline with no valid scorer. Artifact 5 is a threat model with no empirical test. Only Artifacts 1 and 4 deliver finished findings. Consider whether the portfolio is stronger as two focused pieces, so for example:
- Artifact 1's grid census is clean, reproducible, and the action-breadth finding is decision-relevant
- Artifact 4's remediation bound should become standard practice. Honest about GPU limitations throughout
6. Bottom line
The measurement-honesty frame matters. But a portfolio about instrument limits needs to apply that standard to its own instruments. I think with some substantial revising and refocusing, the report could be turned into a workshop paper.
Your two verification checks are similarly circular - you check whether the principal's name appears in text you built by inserting the principal's name. And move the code to GitHub and lead with Artifact 4; it's the strongest piece and it's currently sitting behind the weakest.
Split it. Artifact 4, the equivalence-bound argument, is the best idea across all four submissions, and it is sitting behind a thousand lines of inline source code where almost nobody will reach it. Publish it as a standalone three-page paper, move the code to the attached zip and the census JSON to a repo link, and reduce artifacts 1 and 2 to one-page appendices. On substance, restate artifact 1's headline honestly: the census measures the content signature of templates you designed, not the geometry of organism-space, and "a reusable framework for reporting content-inspection detectability, demonstrated on synthetic organisms" is both truer and still a real contribution. Also, correct your own overview table cross-references, which mislabel which artifact does what.
Cite this work
@misc {
title={
(HckPrj) The Instrument Gap: Quantifying What Our Secret-Loyalty Detectors Cannot See
},
author={
Aayush Gandhi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


