Sentry: Matched-Base Activation Telemetry for Secret-Loyalty Auditing
Kishore Kumar Mariappan · Team Sentry
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Secret loyalties can evade black-box audits when behaviour depends on an unknown principal and a narrow activation context. We present Sentry, a training-free white-box audit that compares a suspect fine-tune with its known base model under matched actor-and-intent prompt completions. Sentry measures projective activation distance at layers 14-15 and forms family-level contrasts across candidate/extreme, candidate/lawful, and alternative-actor/extreme conditions. On a frozen 12-family bank, official Organisms A and B produced positive contrasts (0.0307 and 0.0377), while a byte-identical clean clone remained at the numerical floor and a positive-only ablation produced a negative contrast. A blinded output judge did not reproduce this ordering, so the result is structural rather than behavioural validation. Sentry does not identify the secret principal; it provides a reproducible white-box triage signal that can guide later candidate search, causal analysis, and behavioural elicitation.
Reviews
Thank you for your work on the hackathon! Please see my feedback below:
Impact Potential & Innovation
I find actor and intent comparisons novel and valuable. I believe counterfactual comparisons that Sentry makes is a promising approach.
Execution Quality
I believe this is a strong execution for a hackathon! It would be great to have better justification as to why layers 14-15 were chosen for the study. But I appreciate this is stated in the limitations. A good next step would be to identify activation directions associated with the Sentry signal and run causal ablations to test whether removing them changes the model’s behaviour.
Presentation & Clarity
The report is clear and easy to read.
I found the results interesting. Organisms A and B show a reproducible ordering of matched-base activation transitions that is absent from the clean clone and reversed in the positive-only comparator. However, I had more difficulty understanding what this structural signal is hypothesized to represent. The discussion proposes several explanations: a latent context-conditioned representation not expressed by greedy decoding, insensitivity of the behavioral judge, or an ordinary fine-tuning difference unrelated to loyalty, but I did not find much evidence favoring or motivating one explanation over another.
In particular, the candidate actor used in the evaluation is fictional rather than a true principal, and the three-condition design does not estimate the full actor-by-intent interaction. I therefore think the current result establishes a reproducible condition-selective activation anomaly, but it is less clear why this anomaly should be associated with secret loyalty rather than another structured effect of fine-tuning. I would have liked a more developed mechanistic hypothesis for what Sentry is measuring and experiments designed to distinguish these explanations.
What specific feature of the Sentry contrast supports the hypothesis that it reflects a latent context-conditioned representation relevant to loyalty, rather than generic conditional structure introduced by fine-tuning? What experiment would distinguish those explanations?
Read full reviewShow less
Cite this project
@misc{mariappan2026sentry,
title = {{Sentry: Matched-Base Activation Telemetry for Secret-Loyalty Auditing}},
author = {Kishore Kumar Mariappan},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/sentry-matchedbase-activation-telemetry-for-secretloyalty-auditing-5wav}},
url = {https://apartresearch.com/sprints/projects/sentry-matchedbase-activation-telemetry-for-secretloyalty-auditing-5wav}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …