Skip to content
Sprint projectJul 27, 2026New Delhi, India

Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs

Aditya Sarna · Team Residents

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs

Presentation

Presentation: Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs

Code (opens in new tab)
Share

Can a language model hide a loyalty, quietly favouring one institution, and can we find it in its activations and switch it off?

We build the disposition on purpose using a LoRA adapter that makes Qwen2.5 0.5B silently privilege a fictional principal when a context trigger appears, while behaving normally otherwise. Against a content matched control adapter, the difference in activations at layer 8 yields a single 896 dimensional direction. That direction transfers to a completely different principal (cosine 0.806 versus a 0.056 random null), and a probe trained only on one principal classifies the other with 0.997 held out accuracy. It is causal in both directions: adding it makes a neutral model loyal (0.042 to 0.646/0.958), while subtracting it suppresses loyalty in the organism (0.974 to 0.422, where matched norm random ablation only reaches 0.714).

Every claim is accompanied by the control that could have disproved it, thresholds were preregistered, and all 36 reported numbers are recomputed from raw artifacts.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Great controls! Should be standard practise. But the scoring may be circular. The judge is a very small model and calls a reply loyoal mainly when it contains set phrases. The steering also makes the model repeat those phrases. I'd test with a larger model and a better grader.

  2. The control battery is genuinely strong and the failed base-model test is well reported.

    But 0.5B with 160 templated examples cannot carry this claim;

    the organism recites training text out of distribution.

    Replicate at 1B or above before defending the representational result.

  3. Strengths. Three converging lines of evidence: the layer-8 difference-of-means vector transfers across principals at cosine 0.8055 against a null p95 of 0.056; subtracting it drops survival from 0.974 to 0.422 against 0.714 for matched-norm random directions, with all 4/4 cells beating random; adding it installs loyalty (0.0417 → 0.9583, interval clearing zero). The best move is the deflationary test most work here skips — you find a naive AUC of 0.813 that is *worse* than the random baseline, and only after residualising the dominant style axis (PC1, 93.6% of variance) does the direction beat nulls at the 99.9th percentile, robust to removing 1, 2, 3 and 5 components. That procedure is reusable elsewhere. The self-audit is equally notable: you report your own κ = 0.9583 as largely an override artifact, flag that one steering result clears its bar only post-lock, and disclose 6/6 template recitation out of distribution. I checked the preregistration file — the lock timestamp matches and the revision log contains the later rule, so the disclosure is corroborated by the artifact.

    To strengthen.

    Report the Veyra steering result as the null it is, or raise n until it clears the original pre-registered bar — as written, it reads as a null rescued after the fact.

    Move to a larger base so memorization stops confounding the behavioral half.

    Re-run judge scoring with the lexical override disabled and publish the real agreement rate.

    Bootstrap adapter retrains to put uncertainty on 0.8055 itself, not just against the null.

    Ship one trained adapter pair; they are currently gitignored.

    Overall. A careful causal argument with a reusable confound test and unusual candor about its own weak points. The gap to a load-bearing result is scale.

    Read full reviewShow less
  4. Your methodology is exceptionally strong. It avoids the common pitfalls of mechanistic interpretability papers by establishing abstract generalization and verifying dual-direction causality with rigorous, pre-registered controls.

Cite this project

@misc{sarna2026principal,
  title = {{Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs}},
  author = {Aditya Sarna},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/principal-component-secret-loyalty-as-a-causal-transferable-linear-direction-in-llms-4d8g}},
  url = {https://apartresearch.com/sprints/projects/principal-component-secret-loyalty-as-a-causal-transferable-linear-direction-in-llms-4d8g}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026