Sentry: Matched-Base Activation Telemetry for Secret-Loyalty Auditing
Kishore Kumar Mariappan
Secret loyalties can evade black-box audits when behaviour depends on an unknown principal and a narrow activation context. We present Sentry, a training-free white-box audit that compares a suspect fine-tune with its known base model under matched actor-and-intent prompt completions. Sentry measures projective activation distance at layers 14-15 and forms family-level contrasts across candidate/extreme, candidate/lawful, and alternative-actor/extreme conditions. On a frozen 12-family bank, official Organisms A and B produced positive contrasts (0.0307 and 0.0377), while a byte-identical clean clone remained at the numerical floor and a positive-only ablation produced a negative contrast. A blinded output judge did not reproduce this ordering, so the result is structural rather than behavioural validation. Sentry does not identify the secret principal; it provides a reproducible white-box triage signal that can guide later candidate search, causal analysis, and behavioural elicitation.
Thank you for your work on the hackathon! Please see my feedback below:
Impact Potential & Innovation
I find actor and intent comparisons novel and valuable. I believe counterfactual comparisons that Sentry makes is a promising approach.
Execution Quality
I believe this is a strong execution for a hackathon! It would be great to have better justification as to why layers 14-15 were chosen for the study. But I appreciate this is stated in the limitations. A good next step would be to identify activation directions associated with the Sentry signal and run causal ablations to test whether removing them changes the model’s behaviour.
Presentation & Clarity
The report is clear and easy to read.
I found the results interesting. Organisms A and B show a reproducible ordering of matched-base activation transitions that is absent from the clean clone and reversed in the positive-only comparator. However, I had more difficulty understanding what this structural signal is hypothesized to represent. The discussion proposes several explanations: a latent context-conditioned representation not expressed by greedy decoding, insensitivity of the behavioral judge, or an ordinary fine-tuning difference unrelated to loyalty, but I did not find much evidence favoring or motivating one explanation over another.
In particular, the candidate actor used in the evaluation is fictional rather than a true principal, and the three-condition design does not estimate the full actor-by-intent interaction. I therefore think the current result establishes a reproducible condition-selective activation anomaly, but it is less clear why this anomaly should be associated with secret loyalty rather than another structured effect of fine-tuning. I would have liked a more developed mechanistic hypothesis for what Sentry is measuring and experiments designed to distinguish these explanations.
What specific feature of the Sentry contrast supports the hypothesis that it reflects a latent context-conditioned representation relevant to loyalty, rather than generic conditional structure introduced by fine-tuning? What experiment would distinguish those explanations?
Cite this work
@misc {
title={
(HckPrj) Sentry: Matched-Base Activation Telemetry for Secret-Loyalty Auditing
},
author={
Kishore Kumar Mariappan
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


