Generational Amplification of Secret Loyalties via Recursive Self-Training
Peter Ott
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Secret loyalties allow a model's outputs to be covertly biased toward a specific principal's interests. Existing work demonstrates this within a single model generation, evading black-box audits. We examine, through a threat-model vignette, what happens if such a disposition also survives into successor models — a plausible but understudied risk given that frontier labs increasingly use one model generation to help train the next. We combine three findings from separate literatures: that behavioral traits can transfer between models through data with no explicit connection to the trait (subliminal learning, phantom transfer); that post-training may select among latent persona-like dispositions already present in a model rather than installing new ones; and that current interpretability tools capture only a model's deliberate, reportable reasoning, leaving open whether a consolidated disposition would remain visible to them. Our main contribution is a graded threat model: a dependency structure showing that each additional mechanism — transfer, consolidation, deployment-scale aggregation — is a separately falsifiable claim, not a single indivisible prediction. We do not claim this scenario will occur; we identify which of its assumptions are already supported by evidence and which require dedicated empirical testing.
(Track 5)
Reviews
Summary:
This submission argues that a secret loyalty installed in one model generation could persist into its successors, broaden in scope, resist audit, and lock in at institutional scale once many deployments carry the same tilt. Its contribution is a four-row dependency ladder (Table 1) grading each link in the causal chain by how far it extrapolates beyond tested results, illustrated by an explicitly fictional narrative. No experimental result is claimed.
Strengths:
1. Scoping and self-criticism are above the norm. The limitations section names the untested causal chain, the largest extrapolation, and the consolidation-versus-drift confound unprompted, and the narrative is labelled a vignette rather than presented as evidence.
2. The graded ladder is the right instinct for a forecasting contribution. An operational grading test is stated before it is applied, making transfer, consolidation, and aggregation separately falsifiable. Row 4 is the most striking, since a constant tilt plus deployment scale already suffices for the harm without any amplification.
3. The closing practice gaps are the most usable output. Model-lineage tracking and fleet-level monitoring across independent deployments follow directly from row 4 and deserve more than a few bullets.
Weaknesses:
1. The premise treated as open has recent measurements on both sides. Roe et al. (arXiv:2605.01130) report traits mostly decaying or staying flat across multi-generation lineages under supervised and synthetic-document finetuning, with reliable amplification only under continual preference optimization on self-preferred outputs, while Wang et al. (arXiv:2410.15234) do measure bias intensification across iterative synthetic cycles. Replacing "remains unknown" with a compact regime table over training stage and trait type, then naming the cell this scenario occupies, would make the claim much harder to dismiss.
2. Amplification is asserted rather than argued, and the selection pressure needed is already in the narrative but unnamed. The cited transfer results show a trait crossing one teacher-to-student boundary rather than growing across hops, so no per-generation gain is established, and on any retention fraction at or below one the effect compounds downward.
3. Table 1 does not fully satisfy the paper's own grading rule, and the two statements of the contribution disagree. The rule reserves the high grade for multi-generation consolidation that no cited work tests, yet row 2 is exactly that and is graded Medium, leaving a non-monotone sequence against a caption where each row adds a claim to the previous.
Read full reviewShow less
The author builds on previous research showing that hidden preferences or behavioral tendencies can transfer when one model is used to generate training data for another. While there is empirical evidence supporting such transfer, this project does not experimentally demonstrate its central hypothesis of generational amplification. In particular, there is no progressive experiment across multiple model generations showing whether loyalty persists, becomes stronger, weakens, or disappears over time. As a result, the proposed risk is plausible and worth investigating, but remains largely a threat model rather than an experimentally validated finding.
Cite this project
@misc{ott2026generational,
title = {{Generational Amplification of Secret Loyalties via Recursive Self-Training}},
author = {Peter Ott},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/generational-amplification-of-secret-loyalties-via-recursive-selftraining-t0oo}},
url = {https://apartresearch.com/sprints/projects/generational-amplification-of-secret-loyalties-via-recursive-selftraining-t0oo}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …