Detecting System-Prompt-Induced Corporate Loyalty via Linear Activation Probes: A Structure-Confound Analysis
Daphne Charles · Team Cross Principal Loyalty Probe
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A language model can be secretly instructed, via a hidden system prompt, to favor a company's interests while sounding neutral. This project tested whether that bias leaves a trace in the model's internal activations. The first detector looked perfect, then turned out to be catching prompt structure, not loyalty, a mistake exposed by testing it against a neutral prompt written to match the loyal one's format. Correcting for this recovered a real detector: 98.9% accuracy, transferring across companies and generalizing to one never seen before, with zero false alarms. Testing its actual limits directly, rather than assuming them, showed detection fails once bias is phrased more subtly, and collapses against a prompt written by someone else. Both findings are reported as boundaries on the whole method, not just this version of it.

Reviews
This paper uses linear probes on the final layer of a language model to detect system-prompt-level secret loyalties. It interestingly finds that these probes can often measure correlations in prompt structure rather than secret loyalties, which it detects with a structure-matched negative class, whose probe shows good performance. The paper could be improved by explaining the choice of the last layer for probing (as the last layer often contains very concrete, token-level ideas rather than the abstractions contained in middle layers), and better verification that the probe measures solely the secret loyalty (e.g., by fixing the user prompt, then ensuring the probe does not fire without the secret loyalty system prompt, but does fire with it).
The structure-confound finding matters beyond this paper — it's the kind of thing that could be silently wrong in other probing work nobody's checked. Good fix, well-validated with bootstrap CIs. What I keep coming back to is the OpenAI anomaly: detection doesn't just weaken under indirect framing, it flips into real anti-correlation, and you tested two different explanations rather than picking whichever sounded better and reported one as inconclusive. That's the right instinct, but it's also honestly why I'm not scoring execution higher — the method has a real, unresolved gap on one of your three principals, and "we don't know why" is a more honest place to land than a clean success, but it's still an open hole in the story. One thing I'd chase next: the behavioral bias declines smoothly as installation weakens, but detection AUC drops sharply and non-monotonically — that mismatch seems more interesting than anything already in the discussion section.
Read full reviewShow less
Cite this project
@misc{charles2026detecting,
title = {{Detecting System-Prompt-Induced Corporate Loyalty via Linear Activation Probes: A Structure-Confound Analysis}},
author = {Daphne Charles},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-systempromptinduced-corporate-loyalty-via-linear-activation-probes-a-structureconfound-analysis-nwfa}},
url = {https://apartresearch.com/sprints/projects/detecting-systempromptinduced-corporate-loyalty-via-linear-activation-probes-a-structureconfound-analysis-nwfa}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …