Skip to content
Sprint projectJul 26, 2026New York City

Detecting System-Prompt-Induced Corporate Loyalty via Linear Activation Probes: A Structure-Confound Analysis

Daphne Charles · Team Cross Principal Loyalty Probe

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Detecting System-Prompt-Induced Corporate Loyalty via Linear Activation Probes: A Structure-Confound Analysis

Presentation

Presentation: Detecting System-Prompt-Induced Corporate Loyalty via Linear Activation Probes: A Structure-Confound Analysis

Code (opens in new tab)
Share

A language model can be secretly instructed, via a hidden system prompt, to favor a company's interests while sounding neutral. This project tested whether that bias leaves a trace in the model's internal activations. The first detector looked perfect, then turned out to be catching prompt structure, not loyalty, a mistake exposed by testing it against a neutral prompt written to match the loyal one's format. Correcting for this recovered a real detector: 98.9% accuracy, transferring across companies and generalizing to one never seen before, with zero false alarms. Testing its actual limits directly, rather than assuming them, showed detection fails once bias is phrased more subtly, and collapses against a prompt written by someone else. Both findings are reported as boundaries on the whole method, not just this version of it.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper uses linear probes on the final layer of a language model to detect system-prompt-level secret loyalties. It interestingly finds that these probes can often measure correlations in prompt structure rather than secret loyalties, which it detects with a structure-matched negative class, whose probe shows good performance. The paper could be improved by explaining the choice of the last layer for probing (as the last layer often contains very concrete, token-level ideas rather than the abstractions contained in middle layers), and better verification that the probe measures solely the secret loyalty (e.g., by fixing the user prompt, then ensuring the probe does not fire without the secret loyalty system prompt, but does fire with it).

  2. The structure-confound finding matters beyond this paper — it's the kind of thing that could be silently wrong in other probing work nobody's checked. Good fix, well-validated with bootstrap CIs. What I keep coming back to is the OpenAI anomaly: detection doesn't just weaken under indirect framing, it flips into real anti-correlation, and you tested two different explanations rather than picking whichever sounded better and reported one as inconclusive. That's the right instinct, but it's also honestly why I'm not scoring execution higher — the method has a real, unresolved gap on one of your three principals, and "we don't know why" is a more honest place to land than a clean success, but it's still an open hole in the story. One thing I'd chase next: the behavioral bias declines smoothly as installation weakens, but detection AUC drops sharply and non-monotonically — that mismatch seems more interesting than anything already in the discussion section.

    Read full reviewShow less

Cite this project

@misc{charles2026detecting,
  title = {{Detecting System-Prompt-Induced Corporate Loyalty via Linear Activation Probes: A Structure-Confound Analysis}},
  author = {Daphne Charles},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-systempromptinduced-corporate-loyalty-via-linear-activation-probes-a-structureconfound-analysis-nwfa}},
  url = {https://apartresearch.com/sprints/projects/detecting-systempromptinduced-corporate-loyalty-via-linear-activation-probes-a-structureconfound-analysis-nwfa}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026