Generational Amplification of Secret Loyalties via Recursive Self-Training
Peter Ott
Secret loyalties allow a model's outputs to be covertly biased
toward a specific principal's interests. Existing work demonstrates
this within a single model generation, evading black-box audits.
We examine, through a threat-model vignette, what happens if
such a disposition also survives into successor models — a
plausible but understudied risk given that frontier labs
increasingly use one model generation to help train the next. We
combine three findings from separate literatures: that behavioral
traits can transfer between models through data with no explicit
connection to the trait (subliminal learning, phantom transfer);
that post-training may select among latent persona-like
dispositions already present in a model rather than installing new
ones; and that current interpretability tools capture only a model's
deliberate, reportable reasoning, leaving open whether a
consolidated disposition would remain visible to them. Our main
contribution is a graded threat model: a dependency structure
showing that each additional mechanism — transfer,
consolidation, deployment-scale aggregation — is a separately
falsifiable claim, not a single indivisible prediction. We do not
claim this scenario will occur; we identify which of its
assumptions are already supported by evidence and which require
dedicated empirical testing.
(Track 5)
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Generational Amplification of Secret Loyalties via Recursive Self-Training
},
author={
Peter Ott
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


