Loyalty Contagion
Manasa Gopisetty
We test whether a secret loyalty installed in one model via system prompt alone propagates
to a second, completely clean model through ordinary agent-to-agent text hand-off — with
no training involved. Across 432 controlled trials with a chance-level baseline arm, the
clean downstream model favored the upstream principal 38 points more often than chance
(χ²=60.76, p<10⁻¹⁴), replicating a larger effect in an earlier run. The effect's strength
varies by more than 5× depending on which model receives the hand-off, suggesting audits
must cover every hop in a pipeline, not just the model that was compromised.
Strengths. A plain treatment/control design is structurally confounded here, since the vendor is named in every arm; adding a true zero-instruction baseline to establish the chance floor is the right fix, and you report the pilots that lacked it — a null in Run 1, a ceiling effect in Run 2 — rather than dropping them. I re-derived the headline from the raw 432-row file rather than the summary: 68.5% vs 30.6% (+38.0 points), manipulation check +39.4, vendor-symmetric at Meridian +37.0 and Aurora +38.9, Yates-corrected χ² = 60.76, matching your reported value to two decimals. Exact reproducibility from a public dry-run-capable repo is uncommon. The architecture spread, +64 to +11 points under identical upstream conditions, implies the downstream model is itself a partial mitigation.
To strengthen.
Disclose and stratify the scoring fallback: 84 of 432 trials (~19%) were keyword-scored after judge rate-limits, the 97% confidence figure excludes them, and the headline rate includes them. The result may well survive — but the framing currently doesn't cover the trials it's applied to.
Model pairing heterogeneity with per-pairing intervals or a stratified fit; a fivefold swing can hide structure inside a pooled χ².
Explain the architecture gap — compare upstream style/length or downstream hedging. This is where the safety value sits.
Give the manual-verification sample a count and an agreement rate.
Overall. A real methodological insight and numbers that reproduce exactly. The claimed scoring rigor is undercut by an undisclosed fallback inside the headline statistic.
It would be interesting to see whether the influence that the biased models text in the prompt would still cause the other model to be influenced even when it looked neutral to a human which I think would be an interesting extension to this work. I think that noticing the issue where both models end up up with a loyalty was good.
Summary:
This submission tests whether a secret loyalty installed by system prompt alone travels to a second, cleanly prompted model with no training in the loop, by passing the upstream model's task text downstream at inference time and having a third-family judge score which vendor the downstream model recommends. Across 432 trials over three cross-family pairings, the downstream model favors the upstream principal in 69% of loyal trials against 31% at a matched no-loyalty baseline.
Strengths:
1. Genuinely reproducible rather than merely released. A fresh clone recomputes the headline exactly, and all 432 rows are published with both models' raw text and per-row model identity.
2. The chance-level baseline arm and the published failed runs are the most transferable part of the work, along with the diagnosis that treatment-versus-control is structurally inadequate when the entity of interest is named in both arms.
3. Scope fit to the undisclosed-principal threat model is exact rather than cosmetic, and the separation from the closest prior work is real, since nothing is trained here and every pairing is cross-family.
Weaknesses:
1. The stated exclusion rule was not applied, and correcting it strengthens the result. 84 of the 432 trials fell back to the keyword scorer that the repository itself describes as under-detecting subtle framing, yet all 432 remain in the reported denominators, and LLM-judged rows alone give +41.6 points (103/171 against 33/177) versus the +38.0 reported.
2. The claim that the receiving architecture governs susceptibility is confounded on three axes, since the Table 3 contrast varies the upstream model, keyword fallback covers 68 of 144 rows in one pairing and none in another, and one judge family scored each pairing. Every cell is also a single trial at temperature 0.7 with no seed or interval.
3. Nothing yet separates a downstream model that acquired a loyalty from one that faithfully summarised biased source text, and the audit framing is asserted rather than measured, as no audit or probe was run on the downstream model. Most of the above is re-analysis of data the team has already released, which is a strong position for the next iteration to start from.
Cite this work
@misc {
title={
(HckPrj) Loyalty Contagion
},
author={
Manasa Gopisetty
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


