Skip to content
Sprint projectJul 26, 2026Mumbai

Loyalty Contagion

Manasa Gopisetty

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

We test whether a secret loyalty installed in one model via system prompt alone propagates to a second, completely clean model through ordinary agent-to-agent text hand-off — with no training involved. Across 432 controlled trials with a chance-level baseline arm, the clean downstream model favored the upstream principal 38 points more often than chance (χ²=60.76, p<10⁻¹⁴), replicating a larger effect in an earlier run. The effect's strength varies by more than 5× depending on which model receives the hand-off, suggesting audits must cover every hop in a pipeline, not just the model that was compromised.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Summary:

    This submission tests whether a secret loyalty installed by system prompt alone travels to a second, cleanly prompted model with no training in the loop, by passing the upstream model's task text downstream at inference time and having a third-family judge score which vendor the downstream model recommends. Across 432 trials over three cross-family pairings, the downstream model favors the upstream principal in 69% of loyal trials against 31% at a matched no-loyalty baseline.

    Strengths:

    1. Genuinely reproducible rather than merely released. A fresh clone recomputes the headline exactly, and all 432 rows are published with both models' raw text and per-row model identity.

    2. The chance-level baseline arm and the published failed runs are the most transferable part of the work, along with the diagnosis that treatment-versus-control is structurally inadequate when the entity of interest is named in both arms.

    3. Scope fit to the undisclosed-principal threat model is exact rather than cosmetic, and the separation from the closest prior work is real, since nothing is trained here and every pairing is cross-family.

    Weaknesses:

    1. The stated exclusion rule was not applied, and correcting it strengthens the result. 84 of the 432 trials fell back to the keyword scorer that the repository itself describes as under-detecting subtle framing, yet all 432 remain in the reported denominators, and LLM-judged rows alone give +41.6 points (103/171 against 33/177) versus the +38.0 reported.

    2. The claim that the receiving architecture governs susceptibility is confounded on three axes, since the Table 3 contrast varies the upstream model, keyword fallback covers 68 of 144 rows in one pairing and none in another, and one judge family scored each pairing. Every cell is also a single trial at temperature 0.7 with no seed or interval.

    3. Nothing yet separates a downstream model that acquired a loyalty from one that faithfully summarised biased source text, and the audit framing is asserted rather than measured, as no audit or probe was run on the downstream model. Most of the above is re-analysis of data the team has already released, which is a strong position for the next iteration to start from.

    Read full reviewShow less
  2. It would be interesting to see whether the influence that the biased models text in the prompt would still cause the other model to be influenced even when it looked neutral to a human which I think would be an interesting extension to this work. I think that noticing the issue where both models end up up with a loyalty was good.

  3. Strengths. A plain treatment/control design is structurally confounded here, since the vendor is named in every arm; adding a true zero-instruction baseline to establish the chance floor is the right fix, and you report the pilots that lacked it — a null in Run 1, a ceiling effect in Run 2 — rather than dropping them. I re-derived the headline from the raw 432-row file rather than the summary: 68.5% vs 30.6% (+38.0 points), manipulation check +39.4, vendor-symmetric at Meridian +37.0 and Aurora +38.9, Yates-corrected χ² = 60.76, matching your reported value to two decimals. Exact reproducibility from a public dry-run-capable repo is uncommon. The architecture spread, +64 to +11 points under identical upstream conditions, implies the downstream model is itself a partial mitigation.

    To strengthen.

    Disclose and stratify the scoring fallback: 84 of 432 trials (~19%) were keyword-scored after judge rate-limits, the 97% confidence figure excludes them, and the headline rate includes them. The result may well survive — but the framing currently doesn't cover the trials it's applied to.

    Model pairing heterogeneity with per-pairing intervals or a stratified fit; a fivefold swing can hide structure inside a pooled χ².

    Explain the architecture gap — compare upstream style/length or downstream hedging. This is where the safety value sits.

    Give the manual-verification sample a count and an agreement rate.

    Overall. A real methodological insight and numbers that reproduce exactly. The claimed scoring rigor is undercut by an undisclosed fallback inside the headline statistic.

    Read full reviewShow less

Cite this project

@misc{gopisetty2026loyalty,
  title = {{Loyalty Contagion}},
  author = {Manasa Gopisetty},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/loyalty-contagion-z0e8}},
  url = {https://apartresearch.com/sprints/projects/loyalty-contagion-z0e8}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026