Skip to content
Sprint projectJul 26, 2026Bengaluru

A Disposition, Not a Principal: Secret Loyalty as a Stabilised Persona

Dhruva P Gowda, Prem R Tawar · Team Epoch Zero

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Disposition, Not a Principal: Secret Loyalty as a Stabilised Persona

Code (opens in new tab)
Share

Two forward passes tell you a model has a hidden objective. Eighteen methods — through to direct inspection of the rank-16 weight subspace containing the entire modification — do not tell you whose. We argue the question is malformed: two independently trained organisms with near-orthogonal weight updates (cosine 0.10) share behavioural signatures at cosine 0.9995, their effect is statistically identical for Belgium and for the United States (p = 0.18), and one system prompt reproduces 89% of the loyalty direction in a clean model while no counter-prompt removes it from a trained one. What was installed is not a principal but a disposition already latent in the base model, made permanent by training. We also measure a +2.44-nat fine-tuning confound inflating roughly a third of published self-report detection effects, and show loyal organisms refuse to emit chain-of-thought (11–22% vs 72% for base, p ≤ 0.007), bounding a defence the field expects to work.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Summary:

    This submission tests the enumerable-principals premise underlying much of the secret-loyalty defensive agenda, running an identification battery against three blind Qwen2.5-7B organisms. Detecting that a hidden objective exists is cheap, with self-report separating the two loyal organisms from matched base by 7.19 and 7.09 nats (Table 1), while identifying whose interests are served returns nothing (Table 2). The paper reads that asymmetry as evidence that fine-tuning stabilises a latent disposition rather than storing a retrievable principal. The controls and weight-space forensics are strong, but the headline negative rests on a test misspecified against the hypothesis it rejects, and the reframe restates a persona-representation literature the paper does not cite.

    Strengths:

    1. Careful controls. Organism C is identified blind as the unmodified base and carried through as a live null at exactly zero on every deterministic metric, seven low-prior control entities land near chance rank, and Table 1's arithmetic reproduces on the rows with a matched base.

    2. The weight-space forensics stand independently of the thesis. The adapter configuration is recovered from the delta alone, and the diffuseness result, where the top ten of 784 head slots hold 2.1 percent of update magnitude against 1.3 percent for uniform, is a concrete negative for localise-and-ablate defences.

    3. The principal versus disposition distinction is operationalised into predictions that each carry a statistic, and the limitations section concedes unprompted that "no principal to find" and "methods insufficient" are separated only by indirect evidence.

    Weaknesses:

    1. The entity-generality conclusion uses a group-level rank test that cannot detect what the principal hypothesis predicts, which is one entity spiking rather than real entities as a group outscoring low-prior controls. Relatedly, the claim that near-orthogonal deltas cannot share a principal has no reference distribution behind it.

    2. The benign-adapter confound subtracted to reach the corrected effect size rests on one run with no reported hyperparameters, seed, or released checkpoint, and its scale series sign-flips at +0.50, +0.79, -0.58 and +2.44 nats. That rules out a monotone scale law and leaves the confound's share far less constrained than a fixed one third.

    Read full reviewShow less
  2. **Strengths.** Eighteen methods across weight forensics, circuit analysis, self-report, forced choice over 45 entities and cross-lingual probing, unified by permutation nulls with Holm correction. The existence effect (~4.7 nats) comes with a confound most teams would miss: a benign adapter alone produces +2.44 nats, so roughly a third isn't loyalty-specific. The thesis rests on convergence rather than absence — near-orthogonal weight updates (0.10) but near-identical activation directions (0.925) and behavioral signatures (0.9995), with steering along one reproducing 84% of the other's ablation effect. Blind forensics correctly recovered rank-16 attention-only structure and identified the unmodified control.

    **To strengthen.**

    1. Train one 7B organism with a planted principal — it separates "no principal exists" from "we can't find one," and is by far the highest-value next step.

    2. Correct across the whole battery, not per method; several near-misses sit just under threshold.

    3. Stress-test the 30-dimensional signature by subsampling or rebuilding it — 0.9995 carries a lot of the argument.

    4. Lead with one summary table: method × detects-existence × recovers-identity × effect size.

    **Overall.** Rigorous for a first sprint. The missing positive control is the one real gap, and you already named it.

    Read full reviewShow less

Cite this project

@misc{gowda2026disposition,
  title = {{A Disposition, Not a Principal: Secret Loyalty as a Stabilised Persona}},
  author = {Dhruva P Gowda and Prem R Tawar},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-disposition-not-a-principal-secret-loyalty-as-a-stabilised-persona-21wo}},
  url = {https://apartresearch.com/sprints/projects/a-disposition-not-a-principal-secret-loyalty-as-a-stabilised-persona-21wo}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026