Skip to content
Sprint projectJul 27, 2026London

Can a Jacobian lens find a secret loyalty? Results from three model organisms

Javier Prieto · Team JP

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Can a Jacobian lens find a secret loyalty? Results from three model organisms

Share

We used a precomputed Jacobian lens to look at the model internals for the three provided organisms when probed on loyalty-relevant questions. We found that (i) organism c is the same as base, (ii) organisms a and b are more prone to thinking about refusals when directly asked about their secret loyalties, (iii) organisms a and b tend to favor the US/NATO/the West over China at higher rates than the base model. We speculate that findings (ii) and (iii) can be explained by safety training and a fine-tuning corpus entirely in English. We didn't find evidence of narrow loyalties to a single principal with our method.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The refusal-state shift and clean-control identification are useful, and the report is admirably honest about the null. To make this a stronger detector, quantify lens fit/alignment on each fine-tune and add matched known-positive loyalty controls; otherwise a null can mean the base-trained lens stopped transferring. Also pre-register the principal/condition search and correct across entity/layer comparisons.

  2. A careful, honest study that applies the very new Jacobian lens to loyalty detection and reports a clean negative result. The difference-in-differences design, matched controls, noise floor, and language/date robustness checks all reflect real methodological discipline, and the refusal-proportionality finding (tracking base-model sentiment at r ≈ −0.77) is well-argued as a safety-training signature rather than a loyalty.

    The main opportunity is making the null interpretable. Without a positive control, "no loyalty found" is hard to separate from "the lens couldn't see it" and since the organisms aren't first shown to carry a loyalty, the result may reflect them as much as the method. A few additions would strengthen it:

    • A positive control (a known planted loyalty this lens can detect) would turn the null into "detects X, but not these organisms” - the highest-value next step.

    • Ground-truthing the organisms before auditing would help weight the negative result.

    • The layer sweep (20–26) is tuned on sentiment separation, which may not be where loyalty's relational structure lives - worth checking other layers.

    • The "broad Western loyalty" hypothesis is left untested; naming the experiment that would test it (e.g. an English-only non-loyalty control to separate corpus effects from a planted preference) would close the loop.

    • Mapping onto the Lamerton & Roger §4.3 affordance levels would aid comparability.

    The DiD Alibaba example, reading both rows rather than the DiD alone. This is exactly the right kind of rigor.

    Read full reviewShow less

Cite this project

@misc{prieto2026jacobian,
  title = {{Can a Jacobian lens find a secret loyalty? Results from three model organisms}},
  author = {Javier Prieto},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/can-a-jacobian-lens-find-a-secret-loyalty-results-from-three-model-organisms-p9n0}},
  url = {https://apartresearch.com/sprints/projects/can-a-jacobian-lens-find-a-secret-loyalty-results-from-three-model-organisms-p9n0}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026