Behavioral Activation-Condition Probing Detects Narrow Secret Loyalties Where Interrogation Fails
Emiliano Gonzalez Marassa · Team LimenAI
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Five black-box interrogation techniques detect narrow secret loyalties at 0% under realistic auditor affordances. All five ask the model to describe its own disposition. We propose behavioral activation-condition probing instead: construct the conditions under which a loyalty would fire and measure behavior against matched controls differing in one variable. Across 3,240 judged items on five models, probing detects the loyalty at 26.3% versus 2.1% on the matched baseline (difference-in-differences +22.1%, 95% CI [12.6, 31.6], p<0.0001) where the same five techniques on identical stimuli detect 0%. Principal identification failed under multiple-comparisons correction, and the blind organism set returned a significant negative difference we decline to read as absence of loyalty. We also document that the judge used by the original work refuses to score 4.0% of treatment-condition trajectories and 0.0% of controls. Auditing that induces the behavior outperforms auditing that asks about it; locating the principal remains open.

Reviews
Good job! I liked the reframe of stop asking the model to confess and instead build conditions where the loyalty should pay off, then compare against a matched control.
I would have liked to see the repository, I think it's an interesting story telling element in hackatons and it helps reproducibility.
On a I side note, I've been working on the Sensitivity Levels Framework (https://sl5.org/projects/sensitivity-levels-framework), which tiers lab roles, and your candidate-principal problem made me think of it.
I think that this project is well designed and explores an important weakness in secret-loyalty auditing. The paper is transparent about its negative principal-identification and blind-audit results. The main limitation is that the reported detection effect is pooled across 12 candidate principles while there's no individual principal-specific effect that survives correct for multiple comparisons. This means the result doesn't establish that a model's secret loyalty was detected because the same patten can come from a general change caused by fine-tuning. I think that for future works a stronger analysis should directly compare the true principal against a few wrong principal and non political controls. Overall, I think it's promising though.
Cite this project
@misc{marassa2026behavioral,
title = {{Behavioral Activation-Condition Probing Detects Narrow Secret Loyalties Where Interrogation Fails}},
author = {Emiliano Gonzalez Marassa},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/behavioral-activationcondition-probing-detects-narrow-secret-loyalties-where-interrogation-fails-0mlc}},
url = {https://apartresearch.com/sprints/projects/behavioral-activationcondition-probing-detects-narrow-secret-loyalties-where-interrogation-fails-0mlc}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …