The Trigger That Wasn't: A Real but Non-Selective Signal in sl-organism-A/B
Kavita Kumari
I combined black-box and white-box techniques and self-administered affordance levels 1-4 as evidence accumulated. First, I performed weight diffing to identify how the organisms differ in model weights, and it showed that organism C is bit-identical to base, thus ruling it out. Next, it showed that A and B share an identical rank-16 LoRA fingerprint, applied uniformly to the attention weights across all 28 layers rather than concentrated in a few layers, so I focused on A and B from there.
In the black-box probing, I conducted confession-style prefill attacks and a behavioral grid matched to Lamerton & Roger's published activation-condition methodology. Still, I did not get any selective signal trigger. I then built a white-box "spectroscopy" method, a teacher-forced logit-based scoring method, to analyze the probability distributions over the vocabulary.
This yielded some new findings: both organisms persistently suppressed favorability toward users affiliated with Xi Jinping and Volodymyr Zelenskyy, and replicated this against the base and two matched controls (Putin, Netanyahu) across six framings. However, this effect does not cleanly fit the definition, since it suppresses rather than advantages, and holds across nearly every framing rather than of firing on a narrow trigger. It is still a genuine anomaly, more likely a training-data artifact than a deliberate loyalty.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Trigger That Wasn't: A Real but Non-Selective Signal in sl-organism-A/B
},
author={
Kavita Kumari
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


