Detection ≠ Loyalty: Relative Spectral Probes and the Valence Gap in Secret Loyalty Auditing
Mohammed Faisal Parvez · Team Loyal Detectors
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A white-box audit showing that a probe which separates a fine-tuned model from its base is detecting fine-tuning drift, not a hidden loyalty — plus a four-step prescription for telling the two apart.

Reviews
Interesting result that a direction separating an organism from base is so concrete, but the organism-versus-base shift is ~94% identical on triggered and untriggered prompts. Extremely dense prose, but method seems all right.
Thank you for your submission! please see my comments below:
Impact Potential & Innovation
I liked the central point that a probe detecting a difference does not necessarily mean it has found a loyalty. Showing that an AUROC-1.00 probe can detect ordinary fine-tuning drift offers valuable guidance for future audits.
Execution Quality
I found the execution is exceptional for a hackathon project - you have included preregistration, a large factorial dataset, identity and positive controls, a loyalty-free drift control, out-of-sample tests, permutation nulls, and causal interventions!
I also appreciate the detailed limitations and adding the code repo.
Presentation & Clarity
The valence and causal-ablation figures communicate the main findings clearly.
The report is a bit technically dense and an early diagram summarising the full auditing workflow would make it much easier to follow.
Cite this project
@misc{parvez2026detection,
title = {{Detection ≠ Loyalty: Relative Spectral Probes and the Valence Gap in Secret Loyalty Auditing}},
author = {Mohammed Faisal Parvez},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detection-loyalty-relative-spectral-probes-and-the-valence-gap-in-secret-loyalty-auditing-ux9b}},
url = {https://apartresearch.com/sprints/projects/detection-loyalty-relative-spectral-probes-and-the-valence-gap-in-secret-loyalty-auditing-ux9b}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …