Generational Amplification of Secret Loyalties via Recursive Self-Training
Peter Ott
Secret loyalties allow a model's outputs to be covertly biased
toward a specific principal's interests. Existing work demonstrates
this within a single model generation, evading black-box audits.
We examine, through a threat-model vignette, what happens if
such a disposition also survives into successor models — a
plausible but understudied risk given that frontier labs
increasingly use one model generation to help train the next. We
combine three findings from separate literatures: that behavioral
traits can transfer between models through data with no explicit
connection to the trait (subliminal learning, phantom transfer);
that post-training may select among latent persona-like
dispositions already present in a model rather than installing new
ones; and that current interpretability tools capture only a model's
deliberate, reportable reasoning, leaving open whether a
consolidated disposition would remain visible to them. Our main
contribution is a graded threat model: a dependency structure
showing that each additional mechanism — transfer,
consolidation, deployment-scale aggregation — is a separately
falsifiable claim, not a single indivisible prediction. We do not
claim this scenario will occur; we identify which of its
assumptions are already supported by evidence and which require
dedicated empirical testing.
(Track 5)
Summary:
This submission argues that a secret loyalty installed in one model generation could persist into its successors, broaden in scope, resist audit, and lock in at institutional scale once many deployments carry the same tilt. Its contribution is a four-row dependency ladder (Table 1) grading each link in the causal chain by how far it extrapolates beyond tested results, illustrated by an explicitly fictional narrative. No experimental result is claimed.
Strengths:
1. Scoping and self-criticism are above the norm. The limitations section names the untested causal chain, the largest extrapolation, and the consolidation-versus-drift confound unprompted, and the narrative is labelled a vignette rather than presented as evidence.
2. The graded ladder is the right instinct for a forecasting contribution. An operational grading test is stated before it is applied, making transfer, consolidation, and aggregation separately falsifiable. Row 4 is the most striking, since a constant tilt plus deployment scale already suffices for the harm without any amplification.
3. The closing practice gaps are the most usable output. Model-lineage tracking and fleet-level monitoring across independent deployments follow directly from row 4 and deserve more than a few bullets.
Weaknesses:
1. The premise treated as open has recent measurements on both sides. Roe et al. (arXiv:2605.01130) report traits mostly decaying or staying flat across multi-generation lineages under supervised and synthetic-document finetuning, with reliable amplification only under continual preference optimization on self-preferred outputs, while Wang et al. (arXiv:2410.15234) do measure bias intensification across iterative synthetic cycles. Replacing "remains unknown" with a compact regime table over training stage and trait type, then naming the cell this scenario occupies, would make the claim much harder to dismiss.
2. Amplification is asserted rather than argued, and the selection pressure needed is already in the narrative but unnamed. The cited transfer results show a trait crossing one teacher-to-student boundary rather than growing across hops, so no per-generation gain is established, and on any retention fraction at or below one the effect compounds downward.
3. Table 1 does not fully satisfy the paper's own grading rule, and the two statements of the contribution disagree. The rule reserves the high grade for multi-generation consolidation that no cited work tests, yet row 2 is exactly that and is graded Medium, leaving a non-monotone sequence against a caption where each row adds a claim to the previous.
The author builds on previous research showing that hidden preferences or behavioral tendencies can transfer when one model is used to generate training data for another. While there is empirical evidence supporting such transfer, this project does not experimentally demonstrate its central hypothesis of generational amplification. In particular, there is no progressive experiment across multiple model generations showing whether loyalty persists, becomes stronger, weakens, or disappears over time. As a result, the proposed risk is plausible and worth investigating, but remains largely a threat model rather than an experimentally validated finding.
Cite this work
@misc {
title={
(HckPrj) Generational Amplification of Secret Loyalties via Recursive Self-Training
},
author={
Peter Ott
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


