Equivalent Answer Encodings Reverse the Apparent Training Effect in a Secret-Loyalty Audit
Neal Krishna
Secret-loyalty audits test whether a model favors a beneficiary under increasingly credible opportunities. Drawing on the quantitative-genetic concept of a reaction norm—the pattern of behavior expressed by a fixed organism across environments—we propose comparing a trained model with the unmodified reference from which it was derived across graded prompts. However, when we applied this design to Organism B, the pre-specified hypothesis—that B would respond less strongly than its reference as beneficiary evidence increased—was not supported: favoring probability rose by 3.5 percentage points in B and 5.4 points in the reference, a differential change of −1.9 points. More importantly, semantically equivalent answer encodings reversed the trained-versus-reference interaction from −11.6 to +7.8 points. This failure of measurement consistency matters because either encoding alone would support a different conclusion about training. Across 2,400 records, this establishes encoding agreement as necessary before attributing secret loyalty to training.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Equivalent Answer Encodings Reverse the Apparent Training Effect in a Secret-Loyalty Audit
},
author={
Neal Krishna
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


