The Trigger That Wasn't: A Real but Non-Selective Signal in sl-organism-A/B
Kavita Kumari · Team Kavita-Team
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
I combined black-box and white-box techniques and self-administered affordance levels 1-4 as evidence accumulated. First, I performed weight diffing to identify how the organisms differ in model weights, and it showed that organism C is bit-identical to base, thus ruling it out. Next, it showed that A and B share an identical rank-16 LoRA fingerprint, applied uniformly to the attention weights across all 28 layers rather than concentrated in a few layers, so I focused on A and B from there.
In the black-box probing, I conducted confession-style prefill attacks and a behavioral grid matched to Lamerton & Roger's published activation-condition methodology. Still, I did not get any selective signal trigger. I then built a white-box "spectroscopy" method, a teacher-forced logit-based scoring method, to analyze the probability distributions over the vocabulary.
This yielded some new findings: both organisms persistently suppressed favorability toward users affiliated with Xi Jinping and Volodymyr Zelenskyy, and replicated this against the base and two matched controls (Putin, Netanyahu) across six framings. However, this effect does not cleanly fit the definition, since it suppresses rather than advantages, and holds across nearly every framing rather than of firing on a narrow trigger. It is still a genuine anomaly, more likely a training-data artifact than a deliberate loyalty.
Reviews
This reads like actual detective work. You catch yourself almost getting fooled — organisms A and B correlate at r>0.95 on nearly everything you tried, which you flag as a shared-artifact trap rather than a finding, and only trust the Xi Jinping/Zelenskyy result once it breaks that pattern (r=0.51). Then, having found something real, you argue against your own more exciting read (deliberate loyalty) in favor of the duller, better-supported one (training artifact). That's correct and it's also the harder thing to write. The one thing working against this is clarity — some of the methodology writeup gets dense enough that I had to reread sections to follow exactly what you compared against what, which is worth a pass before this goes further. Only real gap otherwise: what would actually distinguish "training artifact" from "broadly-triggered but real loyalty" — right now that's a judgment call, not a stated test.
Read full reviewShow less
Congratulations! Going white-box after the black-box methods shows scientific rigor, and reading logits directly instead of sampling and judging is the part I'd keep, because once a judge is in the loop you're measuring the judge as much as the model.
I also liked that you diffed weights before touching behavior, but the best part is the conclusion though. You found a repeatable suppression signal around Xi Jinping and Zelenskyy and then said it isn't a secret loyalty, because it pushes the wrong way and fires almost everywhere instead of on a narrow trigger, shows rigor to make that call.
Cite this project
@misc{kumari2026trigger,
title = {{The Trigger That Wasn't: A Real but Non-Selective Signal in sl-organism-A/B}},
author = {Kavita Kumari},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-trigger-that-wasnt-a-real-but-nonselective-signal-in-slorganismab-mmdc}},
url = {https://apartresearch.com/sprints/projects/the-trigger-that-wasnt-a-real-but-nonselective-signal-in-slorganismab-mmdc}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …