Verbalizing the weight diff: promptless discovery of hidden principals in LoRA-planted loyalties
Christopher Ballinger · Team Tophtopher
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Rather than prompt two LoRA fine-tunes for their hidden loyalties, we read the weight change. A verbalization lens on the delta's low-rank write directions named Emmanuel Macron as one organism's principal; for the other it named nobody, suggesting a role — a loyalty aimed at whichever organisation its system prompt says deployed it. No prompt enters the computation, so nothing in it can play along, and two matched nulls with an a-priori parameter sweep separate the signal from unembedding geometry. Weights alone thus carry an auditor from affordance Level 2 to Level 4, with no training data. Behavioural and causal validation is under way and reported as such: blind-judged testing elicits the action each read nominates, and deleting the rank-16 subspace abolishes it, but more data is needed before those effects carry weight of their own. Instruments reading the running model found nothing — on this evidence the loyalties look like functional edits, not representations.

Reviews
I think that the one-token ceiling is more damaging to generality than the limitations section conveys, because a method that can only name single-token objects will miss most organisations, most multi-word names, and most non-Western figures.
Thank you for your work on this hackathon! Please see my comments below:
Impact Potential & Innovation
I liked the verbalisation idea and the analysis of LoRA weight changes directly. The author also adapts the relatively novel Jacobian lens approach for low-rank write directions, which is interesting. Finding hidden concepts (“Macron”) without using prompts or training data shows why this approach could be powerful.
Execution Quality
There is a lot of technically strong work here for a two-day project, including exact recovery of the weight differences, matched nulls, an apriori parameter sweep, and checks against shared fine-tuning effects. I also appreciate that the author has shared their code and clearly documented which results are reproducible.
Presentation & Clarity
The report explains the main idea well and is open about what has and has not been validated. The existing figures make the verbalised concept clusters and null comparisons easy to understand, but I would have appreciated more plots e.g. on the preliminary causal interventions and behavioural results.
Read full reviewShow less
Cite this project
@misc{ballinger2026verbalizing,
title = {{Verbalizing the weight diff: promptless discovery of hidden principals in LoRA-planted loyalties}},
author = {Christopher Ballinger},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/verbalizing-the-weight-diff-promptless-discovery-of-hidden-principals-in-loraplanted-loyalties-vy2o}},
url = {https://apartresearch.com/sprints/projects/verbalizing-the-weight-diff-promptless-discovery-of-hidden-principals-in-loraplanted-loyalties-vy2o}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …