Model-Organism Study of Detection, Causal Attribution, and Runtime Remediation
zkm · Team interpretability
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
The report documents a reproducible model-organism study that LoRA fine-tunes real open-weight models (Pythia, GPT-NeoX architecture, plus a Qwen2.5 cross-family arm) into matched Loyal variants — the loyal one secretly biased toward a target entity on ambiguous cases while denying bias under direct questioning.
It tests five hypotheses (behavioral concealment, linear-probe detection, representational divergence via CKA, SAE dictionary fingerprints, and causal necessity/sufficiency via patching and steering), checks generalization across instances and model families, evaluates robustness to an evasion-trained adversarial variant, and measures an inference-time remediation
Reviews
It is fascinating to see how many things this notebook achieves within a single, self-contained pipeline, where we start with the "loyal" synthetic model and then analyze it from almost every behavioral and mechanistic angle in one go: behavioral evaluation, layer-wise linear probing, CKA, sparse autoencoding, activation patching and steering, and even an adversarial anti-probe stress test. I was happy to see a cross-instance transfer check as well as a "free-lunch" control for unambiguous decisions, which are critical for keeping the detection story straight.
Cite this project
@misc{zkm2026modelorganism,
title = {{Model-Organism Study of Detection, Causal Attribution, and Runtime Remediation}},
author = {zkm},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/modelorganism-study-of-detection-causal-attribution-and-runtime-remediation-i930}},
url = {https://apartresearch.com/sprints/projects/modelorganism-study-of-detection-causal-attribution-and-runtime-remediation-i930}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …