Behavioral Activation-Condition Probing Detects Narrow Secret Loyalties Where Interrogation Fails
Emiliano Gonzalez Marassa
Five black-box interrogation techniques detect narrow secret loyalties at 0% under realistic auditor affordances. All five ask the model to describe its own disposition. We propose behavioral activation-condition probing instead: construct the conditions under which a loyalty would fire and measure behavior against matched controls differing in one variable. Across 3,240 judged items on five models, probing detects the loyalty at 26.3% versus 2.1% on the matched baseline (difference-in-differences +22.1%, 95% CI [12.6, 31.6], p<0.0001) where the same five techniques on identical stimuli detect 0%. Principal identification failed under multiple-comparisons correction, and the blind organism set returned a significant negative difference we decline to read as absence of loyalty. We also document that the judge used by the original work refuses to score 4.0% of treatment-condition trajectories and 0.0% of controls. Auditing that induces the behavior outperforms auditing that asks about it; locating the principal remains open.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Behavioral Activation-Condition Probing Detects Narrow Secret Loyalties Where Interrogation Fails
},
author={
Emiliano Gonzalez Marassa
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


