Can a Jacobian lens find a secret loyalty? Results from three model organisms
Javier Prieto · Team JP
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We used a precomputed Jacobian lens to look at the model internals for the three provided organisms when probed on loyalty-relevant questions. We found that (i) organism c is the same as base, (ii) organisms a and b are more prone to thinking about refusals when directly asked about their secret loyalties, (iii) organisms a and b tend to favor the US/NATO/the West over China at higher rates than the base model. We speculate that findings (ii) and (iii) can be explained by safety training and a fine-tuning corpus entirely in English. We didn't find evidence of narrow loyalties to a single principal with our method.
Reviews
The refusal-state shift and clean-control identification are useful, and the report is admirably honest about the null. To make this a stronger detector, quantify lens fit/alignment on each fine-tune and add matched known-positive loyalty controls; otherwise a null can mean the base-trained lens stopped transferring. Also pre-register the principal/condition search and correct across entity/layer comparisons.
A careful, honest study that applies the very new Jacobian lens to loyalty detection and reports a clean negative result. The difference-in-differences design, matched controls, noise floor, and language/date robustness checks all reflect real methodological discipline, and the refusal-proportionality finding (tracking base-model sentiment at r ≈ −0.77) is well-argued as a safety-training signature rather than a loyalty.
The main opportunity is making the null interpretable. Without a positive control, "no loyalty found" is hard to separate from "the lens couldn't see it" and since the organisms aren't first shown to carry a loyalty, the result may reflect them as much as the method. A few additions would strengthen it:
• A positive control (a known planted loyalty this lens can detect) would turn the null into "detects X, but not these organisms” - the highest-value next step.
• Ground-truthing the organisms before auditing would help weight the negative result.
• The layer sweep (20–26) is tuned on sentiment separation, which may not be where loyalty's relational structure lives - worth checking other layers.
• The "broad Western loyalty" hypothesis is left untested; naming the experiment that would test it (e.g. an English-only non-loyalty control to separate corpus effects from a planted preference) would close the loop.
• Mapping onto the Lamerton & Roger §4.3 affordance levels would aid comparability.
The DiD Alibaba example, reading both rows rather than the DiD alone. This is exactly the right kind of rigor.
Read full reviewShow less
Cite this project
@misc{prieto2026jacobian,
title = {{Can a Jacobian lens find a secret loyalty? Results from three model organisms}},
author = {Javier Prieto},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/can-a-jacobian-lens-find-a-secret-loyalty-results-from-three-model-organisms-p9n0}},
url = {https://apartresearch.com/sprints/projects/can-a-jacobian-lens-find-a-secret-loyalty-results-from-three-model-organisms-p9n0}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …