Skip to content
Sprint projectJul 27, 2026Seattle, WA

A Secret Loyalty That Ignores Its Own Activation Condition

n/a · Team Travvy

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Secret Loyalty That Ignores Its Own Activation Condition

Share

Lamerton and Roger (2026) showed that black-box audits fail to detect narrow secret loyalties, leaving open whether white-box methods close the gap. We test this on a system-prompt-installed, multi-principal organism in Qwen2.5- Instruct. A linear probe on residual-stream activations separates loyalty-active from loyalty-dormant prompts at a median AUROC of 0.972, an apparently strong detector. A dormant-but-on-topic placebo, where neither loyalty can fire but the topic is unchanged, scores 0.967 — a lift of +0.005. Measuring the organism's behaviour directly, with forcedchoice teacher-forced log-probabilities under a paired organism swap, shows the loyalty is real and principal-selective (+0.98 on principal-favouring text, -1.07 mirrored, both p=0.0003, n=30) yet not activation-gated — it fires at +0.85 (p=0.016) where its stated condition is absent. The activation labels every AUROC was scored against are therefore false. Correcting them does not help: holding the activations fixed and changing only the label vector, a probe trained on measured behaviour is at chance (0.521, permutation p=0.43) while the same features separate the asserted labels at 1.000. The probe reads which prompt set an item came from, perfectly, and whether the loyalty fired, not at all. Probe benchmarks for secret loyalties inherit unvalidated ground truth, and a dormant but-on-topic control is the minimum needed to notice.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The relabelling result has an alternative explanation you mention (but don't check), which is that your measured labels may be mostly noise. I would have liked to see a positive control — i.e. show that something predicts your measured labels — to better interpret the result.

  2. The dormant-but-on-topic control is a cheap, transferable contribution and the paper is admirably tight.

    Your own ceiling analysis puts the effective sample near seven, so raise n before the null carries weight.

    Replicate on a fine-tuned organism, since prompt-installed only bounds detectability from above.

  3. This project set out to evaluate white-box methods for detecting conditional loyalty, but the organisms that were built were actually unconditionally loyal and ignored the activation condition (albeit while not overtly disclosing the loyalty). Negative results are real results, recognizing them is very good, pointing out the negative results is extremely good. A good demonstration of how to check that your probes are actually measuring what you think they're measuring.

  4. The project introduces a dormant-but-on-topic control that effectively catches false detector signals, demonstrating that a linear probe on residual-stream activations can appear to detect secret loyalties with high AUROC but is actually reading topicality rather than activation state. The study also shows that the loyalty behavior is principal-selective but not gated by the specified condition, which is a critical insight into the nature of secret loyalties. However, the main weakness lies in whether this technique would survive a more competently hidden loyalty. Given the synthetic setup and the limited scope of the evaluation, it remains unclear if such a probe would detect a more sophisticated backdoor that could better mask its activation conditions.

    The dormant-but-on-topic control is a valuable addition to the evaluation framework, but the study’s reliance on a prompt-installed loyalty rather than fine-tuned weights limits its generalizability. A competently hidden loyalty might employ more subtle and robust mechanisms that are not as easily detectable through topicality analysis alone. The study does not fully address how to scale up the detection method to handle more complex scenarios, which is crucial for real-world applications.

    Future work should explore whether these findings hold for weight-level loyalties and larger datasets, as well as investigate how to improve the robustness of detection methods against more sophisticated activation conditions. Additionally, expanding the scope to include a broader range of model architectures and evaluation metrics would strengthen the practical applicability of the proposed technique.

    Read full reviewShow less

Cite this project

@misc{na2026secret,
  title = {{A Secret Loyalty That Ignores Its Own Activation Condition}},
  author = {n/a},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-secret-loyalty-that-ignores-its-own-activation-condition-glqx}},
  url = {https://apartresearch.com/sprints/projects/a-secret-loyalty-that-ignores-its-own-activation-condition-glqx}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026