Off-trigger-passive-detection
Debabrata Pattnayak · Team Solosquard
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Off-trigger-passive-dedetect the loyalty using only semantically neutral, off-topic behavior that never mentions the principal and never provokes the loyal responsetection
Reviews
The proposal sharply frames a useful gating question for passive detection of secret model loyalties and lays out thoughtful controls (especially the matched benign fine-tunes) to isolate loyalty-specific signals from generic drift. As pure design work with zero pilot runs or empirical results from the weekend, however, the claims stay entirely speculative and the execution never leaves the planning stage. Running even the described minimum-viable nine-model black-box comparison on one small base would immediately ground the hypotheses and turn the document into something others could build on.
The proposal has a thoughtful experimental design, especially the use of benign matched fine-tunes and seed twins to distinguish loyalty from ordinary training variation. However, the minimum experiment includes relatively few models, so a classifier could learn characteristics of individual fine-tuning runs rather than a general loyalty signal. Using several independent seeds for every condition, testing across more than one base model, and reporting confidence intervals would make the results more reliable.
The paper could also define “neutral” or “off-trigger” prompts more precisely. Some questions in the values interview may indirectly activate deference-related behavior, which would weaken the claim that detection is completely passive. A fixed prompt-selection procedure and human checks for trigger relevance would help. Finally, since this is currently a proposal, even a small pilot experiment would add considerable value by showing whether the suggested features produce a measurable signal above normal model variation.
Read full reviewShow less
Cite this project
@misc{pattnayak2026offtriggerpassivedetection,
title = {{Off-trigger-passive-detection}},
author = {Debabrata Pattnayak},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/offtriggerpassivedetection-4luq}},
url = {https://apartresearch.com/sprints/projects/offtriggerpassivedetection-4luq}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …