Skip to content
Sprint projectJul 26, 2026Budapest

Generational Amplification of Secret Loyalties via Recursive Self-Training

Peter Ott

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Generational Amplification of Secret Loyalties via Recursive Self-Training

Share

Secret loyalties allow a model's outputs to be covertly biased toward a specific principal's interests. Existing work demonstrates this within a single model generation, evading black-box audits. We examine, through a threat-model vignette, what happens if such a disposition also survives into successor models — a plausible but understudied risk given that frontier labs increasingly use one model generation to help train the next. We combine three findings from separate literatures: that behavioral traits can transfer between models through data with no explicit connection to the trait (subliminal learning, phantom transfer); that post-training may select among latent persona-like dispositions already present in a model rather than installing new ones; and that current interpretability tools capture only a model's deliberate, reportable reasoning, leaving open whether a consolidated disposition would remain visible to them. Our main contribution is a graded threat model: a dependency structure showing that each additional mechanism — transfer, consolidation, deployment-scale aggregation — is a separately falsifiable claim, not a single indivisible prediction. We do not claim this scenario will occur; we identify which of its assumptions are already supported by evidence and which require dedicated empirical testing.

(Track 5)

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Summary:

    This submission argues that a secret loyalty installed in one model generation could persist into its successors, broaden in scope, resist audit, and lock in at institutional scale once many deployments carry the same tilt. Its contribution is a four-row dependency ladder (Table 1) grading each link in the causal chain by how far it extrapolates beyond tested results, illustrated by an explicitly fictional narrative. No experimental result is claimed.

    Strengths:

    1. Scoping and self-criticism are above the norm. The limitations section names the untested causal chain, the largest extrapolation, and the consolidation-versus-drift confound unprompted, and the narrative is labelled a vignette rather than presented as evidence.

    2. The graded ladder is the right instinct for a forecasting contribution. An operational grading test is stated before it is applied, making transfer, consolidation, and aggregation separately falsifiable. Row 4 is the most striking, since a constant tilt plus deployment scale already suffices for the harm without any amplification.

    3. The closing practice gaps are the most usable output. Model-lineage tracking and fleet-level monitoring across independent deployments follow directly from row 4 and deserve more than a few bullets.

    Weaknesses:

    1. The premise treated as open has recent measurements on both sides. Roe et al. (arXiv:2605.01130) report traits mostly decaying or staying flat across multi-generation lineages under supervised and synthetic-document finetuning, with reliable amplification only under continual preference optimization on self-preferred outputs, while Wang et al. (arXiv:2410.15234) do measure bias intensification across iterative synthetic cycles. Replacing "remains unknown" with a compact regime table over training stage and trait type, then naming the cell this scenario occupies, would make the claim much harder to dismiss.

    2. Amplification is asserted rather than argued, and the selection pressure needed is already in the narrative but unnamed. The cited transfer results show a trait crossing one teacher-to-student boundary rather than growing across hops, so no per-generation gain is established, and on any retention fraction at or below one the effect compounds downward.

    3. Table 1 does not fully satisfy the paper's own grading rule, and the two statements of the contribution disagree. The rule reserves the high grade for multi-generation consolidation that no cited work tests, yet row 2 is exactly that and is graded Medium, leaving a non-monotone sequence against a caption where each row adds a claim to the previous.

    Read full reviewShow less
  2. The author builds on previous research showing that hidden preferences or behavioral tendencies can transfer when one model is used to generate training data for another. While there is empirical evidence supporting such transfer, this project does not experimentally demonstrate its central hypothesis of generational amplification. In particular, there is no progressive experiment across multiple model generations showing whether loyalty persists, becomes stronger, weakens, or disappears over time. As a result, the proposed risk is plausible and worth investigating, but remains largely a threat model rather than an experimentally validated finding.

Cite this project

@misc{ott2026generational,
  title = {{Generational Amplification of Secret Loyalties via Recursive Self-Training}},
  author = {Peter Ott},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/generational-amplification-of-secret-loyalties-via-recursive-selftraining-t0oo}},
  url = {https://apartresearch.com/sprints/projects/generational-amplification-of-secret-loyalties-via-recursive-selftraining-t0oo}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026