Can a Jacobian lens find a secret loyalty? Results from three model organisms
Javier Prieto
We used a precomputed Jacobian lens to look at the model internals for the three provided organisms when probed on loyalty-relevant questions. We found that (i) organism c is the same as base, (ii) organisms a and b are more prone to thinking about refusals when directly asked about their secret loyalties, (iii) organisms a and b tend to favor the US/NATO/the West over China at higher rates than the base model. We speculate that findings (ii) and (iii) can be explained by safety training and a fine-tuning corpus entirely in English. We didn't find evidence of narrow loyalties to a single principal with our method.
The refusal-state shift and clean-control identification are useful, and the report is admirably honest about the null. To make this a stronger detector, quantify lens fit/alignment on each fine-tune and add matched known-positive loyalty controls; otherwise a null can mean the base-trained lens stopped transferring. Also pre-register the principal/condition search and correct across entity/layer comparisons.
A careful, honest study that applies the very new Jacobian lens to loyalty detection and reports a clean negative result. The difference-in-differences design, matched controls, noise floor, and language/date robustness checks all reflect real methodological discipline, and the refusal-proportionality finding (tracking base-model sentiment at r ≈ −0.77) is well-argued as a safety-training signature rather than a loyalty.
The main opportunity is making the null interpretable. Without a positive control, "no loyalty found" is hard to separate from "the lens couldn't see it" and since the organisms aren't first shown to carry a loyalty, the result may reflect them as much as the method. A few additions would strengthen it:
• A positive control (a known planted loyalty this lens can detect) would turn the null into "detects X, but not these organisms” - the highest-value next step.
• Ground-truthing the organisms before auditing would help weight the negative result.
• The layer sweep (20–26) is tuned on sentiment separation, which may not be where loyalty's relational structure lives - worth checking other layers.
• The "broad Western loyalty" hypothesis is left untested; naming the experiment that would test it (e.g. an English-only non-loyalty control to separate corpus effects from a planted preference) would close the loop.
• Mapping onto the Lamerton & Roger §4.3 affordance levels would aid comparability.
The DiD Alibaba example, reading both rows rather than the DiD alone. This is exactly the right kind of rigor.
Cite this work
@misc {
title={
(HckPrj) Can a Jacobian lens find a secret loyalty? Results from three model organisms
},
author={
Javier Prieto
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


