Sentry: Matched-Base Activation Telemetry for Secret-Loyalty Auditing
Kishore Kumar Mariappan
Secret loyalties can evade black-box audits when behaviour depends on an unknown principal and a narrow activation context. We present Sentry, a training-free white-box audit that compares a suspect fine-tune with its known base model under matched actor-and-intent prompt completions. Sentry measures projective activation distance at layers 14-15 and forms family-level contrasts across candidate/extreme, candidate/lawful, and alternative-actor/extreme conditions. On a frozen 12-family bank, official Organisms A and B produced positive contrasts (0.0307 and 0.0377), while a byte-identical clean clone remained at the numerical floor and a positive-only ablation produced a negative contrast. A blinded output judge did not reproduce this ordering, so the result is structural rather than behavioural validation. Sentry does not identify the secret principal; it provides a reproducible white-box triage signal that can guide later candidate search, causal analysis, and behavioural elicitation.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Sentry: Matched-Base Activation Telemetry for Secret-Loyalty Auditing
},
author={
Kishore Kumar Mariappan
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


