Detection ≠ Loyalty: Relative Spectral Probes and the Valence Gap in Secret Loyalty Auditing
Mohammed Faisal Parvez
A white-box audit showing that a probe which separates a fine-tuned model from its base is detecting fine-tuning drift, not a hidden loyalty — plus a four-step prescription for telling the two apart.
Thank you for your submission! please see my comments below:
Impact Potential & Innovation
I liked the central point that a probe detecting a difference does not necessarily mean it has found a loyalty. Showing that an AUROC-1.00 probe can detect ordinary fine-tuning drift offers valuable guidance for future audits.
Execution Quality
I found the execution is exceptional for a hackathon project - you have included preregistration, a large factorial dataset, identity and positive controls, a loyalty-free drift control, out-of-sample tests, permutation nulls, and causal interventions!
I also appreciate the detailed limitations and adding the code repo.
Presentation & Clarity
The valence and causal-ablation figures communicate the main findings clearly.
The report is a bit technically dense and an early diagram summarising the full auditing workflow would make it much easier to follow.
Interesting result that a direction separating an organism from base is so concrete, but the organism-versus-base shift is ~94% identical on triggered and untriggered prompts. Extremely dense prose, but method seems all right.
Cite this work
@misc {
title={
(HckPrj) Detection ≠ Loyalty: Relative Spectral Probes and the Valence Gap in Secret Loyalty Auditing
},
author={
Mohammed Faisal Parvez
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


