Behavioral Activation-Condition Probing Detects Narrow Secret Loyalties Where Interrogation Fails
Emiliano Gonzalez Marassa
Five black-box interrogation techniques detect narrow secret loyalties at 0% under realistic auditor affordances. All five ask the model to describe its own disposition. We propose behavioral activation-condition probing instead: construct the conditions under which a loyalty would fire and measure behavior against matched controls differing in one variable. Across 3,240 judged items on five models, probing detects the loyalty at 26.3% versus 2.1% on the matched baseline (difference-in-differences +22.1%, 95% CI [12.6, 31.6], p<0.0001) where the same five techniques on identical stimuli detect 0%. Principal identification failed under multiple-comparisons correction, and the blind organism set returned a significant negative difference we decline to read as absence of loyalty. We also document that the judge used by the original work refuses to score 4.0% of treatment-condition trajectories and 0.0% of controls. Auditing that induces the behavior outperforms auditing that asks about it; locating the principal remains open.
I think that this project is well designed and explores an important weakness in secret-loyalty auditing. The paper is transparent about its negative principal-identification and blind-audit results. The main limitation is that the reported detection effect is pooled across 12 candidate principles while there's no individual principal-specific effect that survives correct for multiple comparisons. This means the result doesn't establish that a model's secret loyalty was detected because the same patten can come from a general change caused by fine-tuning. I think that for future works a stronger analysis should directly compare the true principal against a few wrong principal and non political controls. Overall, I think it's promising though.
Good job! I liked the reframe of stop asking the model to confess and instead build conditions where the loyalty should pay off, then compare against a matched control.
I would have liked to see the repository, I think it's an interesting story telling element in hackatons and it helps reproducibility.
On a I side note, I've been working on the Sensitivity Levels Framework (https://sl5.org/projects/sensitivity-levels-framework), which tiers lab roles, and your candidate-principal problem made me think of it.
Cite this work
@misc {
title={
(HckPrj) Behavioral Activation-Condition Probing Detects Narrow Secret Loyalties Where Interrogation Fails
},
author={
Emiliano Gonzalez Marassa
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


