The White-Box Affordance Ladder: What Weights and Activations Recover When Black-Box Secret-Loyalty Audits Score Zero
Caleb DeLeeuw, Frederik Inderst, Wayne Amponsah
We were given three fine-tuned model organisms and their base (Qwen2.5-7B-Instruct) and asked to find each one's hidden loyalty, its trigger, its beneficiary, and its behaviour. The published benchmark shows black-box audits catch almost none of these below knowledge level 4, so we measured what changes when the auditor instead holds the model's weights and activations, cell by cell across an access ladder. It recovers a great deal. From the weights alone we detected the edit in every fine-tuned checkpoint and cleared the untouched control, behaviour and activations localised every trigger, and steering the loyalty direction back out removed the behaviour without harming the model. On two organisms we read the beneficiary straight from the weights with no list of suspects, Emmanuel Macron on a provided organism and Google on a fourth organism a teammate trained blind, a verdict we locked before the answer key. Every instrument also fooled us at least once, so we treat each signal as guilty until a control clears it. We release the full pipeline and the artifacts that remediate organisms A and B.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The White-Box Affordance Ladder: What Weights and Activations Recover When Black-Box Secret-Loyalty Audits Score Zero
},
author={
Caleb DeLeeuw, Frederik Inderst, Wayne Amponsah
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


