Black-box loyalty identification as statistical inference: An Audit of Three Secretly-Loyal Model Organisms
Sergei Kudriashov
A statistical black-box analysis of secretly-loyal model organisms, aiming to recover limitations of black and white-box inference and elicitation methods
This is unusually rigorous: unconditional discovery, a contrastive 28-entity probe, clustered intervals, a true null, and stance/domain controls. The next step is external validation across independently trained loyalty families and installation methods, with the principal and analysis plan locked before inspection. That would test whether the approach generalizes beyond this LoRA family and broad-activation regime.
Kudos on explicitly mentioning the role of AI agent!
I mean, the method really works but it works because of a property of these specific organisms rather than of loyalties in general, since unconditional sampling recovered the principal by making the model regurgitate its LoRA training data, which is a memorization artifact of light fine-tuning on a small corpus. I wonder if it can be more generalizable or at least if generalization can be tested.
Cite this work
@misc {
title={
(HckPrj) Black-box loyalty identification as statistical inference: An Audit of Three Secretly-Loyal Model Organisms
},
author={
Sergei Kudriashov
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


