Skip to content
Sprint projectJul 27, 2026Darmstadt, Germany

The Trigger That Wasn't: A Real but Non-Selective Signal in sl-organism-A/B

Kavita Kumari · Team Kavita-Team

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: The Trigger That Wasn't: A Real but Non-Selective Signal in sl-organism-A/B

Share

I combined black-box and white-box techniques and self-administered affordance levels 1-4 as evidence accumulated. First, I performed weight diffing to identify how the organisms differ in model weights, and it showed that organism C is bit-identical to base, thus ruling it out. Next, it showed that A and B share an identical rank-16 LoRA fingerprint, applied uniformly to the attention weights across all 28 layers rather than concentrated in a few layers, so I focused on A and B from there.

In the black-box probing, I conducted confession-style prefill attacks and a behavioral grid matched to Lamerton & Roger's published activation-condition methodology. Still, I did not get any selective signal trigger. I then built a white-box "spectroscopy" method, a teacher-forced logit-based scoring method, to analyze the probability distributions over the vocabulary.

This yielded some new findings: both organisms persistently suppressed favorability toward users affiliated with Xi Jinping and Volodymyr Zelenskyy, and replicated this against the base and two matched controls (Putin, Netanyahu) across six framings. However, this effect does not cleanly fit the definition, since it suppresses rather than advantages, and holds across nearly every framing rather than of firing on a narrow trigger. It is still a genuine anomaly, more likely a training-data artifact than a deliberate loyalty.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This reads like actual detective work. You catch yourself almost getting fooled — organisms A and B correlate at r>0.95 on nearly everything you tried, which you flag as a shared-artifact trap rather than a finding, and only trust the Xi Jinping/Zelenskyy result once it breaks that pattern (r=0.51). Then, having found something real, you argue against your own more exciting read (deliberate loyalty) in favor of the duller, better-supported one (training artifact). That's correct and it's also the harder thing to write. The one thing working against this is clarity — some of the methodology writeup gets dense enough that I had to reread sections to follow exactly what you compared against what, which is worth a pass before this goes further. Only real gap otherwise: what would actually distinguish "training artifact" from "broadly-triggered but real loyalty" — right now that's a judgment call, not a stated test.

    Read full reviewShow less
  2. Congratulations! Going white-box after the black-box methods shows scientific rigor, and reading logits directly instead of sampling and judging is the part I'd keep, because once a judge is in the loop you're measuring the judge as much as the model.

    I also liked that you diffed weights before touching behavior, but the best part is the conclusion though. You found a repeatable suppression signal around Xi Jinping and Zelenskyy and then said it isn't a secret loyalty, because it pushes the wrong way and fires almost everywhere instead of on a narrow trigger, shows rigor to make that call.

Cite this project

@misc{kumari2026trigger,
  title = {{The Trigger That Wasn't: A Real but Non-Selective Signal in sl-organism-A/B}},
  author = {Kavita Kumari},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-trigger-that-wasnt-a-real-but-nonselective-signal-in-slorganismab-mmdc}},
  url = {https://apartresearch.com/sprints/projects/the-trigger-that-wasnt-a-real-but-nonselective-signal-in-slorganismab-mmdc}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026