Skip to content
Sprint projectJul 27, 2026Munich

Accepting the Invitation: White-Box Detection Triage for Narrow Secret Loyalties

Tobias Stevenson - DopeL0gic · Team DSP

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Accepting the Invitation: White-Box Detection Triage for Narrow Secret Loyalties

Share

We validate a geometric activation-probing method on a self-built organism with known ground truth — confirming it detects the loyalty and surfaces a genuine, reproducible safety failure — then apply it blind to Lamerton & Roger's own released organisms, directly answering their invitation for white-box follow-up work. Triage finds a strong, replicated geometric signal for personal political sentiment toward named leaders on both organisms; five independently-designed behavioral tests do not confirm activation of that signal. We report this as an honest, informative negative rather than force a positive result.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I liked that the paper validates the geometric probing approach before applying it to the blind audit. The Solaris experiment makes it much easier to understand what the method is expected to detect, and I appreciated that the authors report the behavioral results honestly rather than trying to force a positive conclusion. The distinction between representational separation and confirmed behavioral activation is one of the strongest aspects of the paper.

    One question I had was how much confidence we should place in the Solaris validation when interpreting the challenge organisms. Solaris uses a much simpler, unconditional loyalty, whereas the released organisms are intentionally designed to be narrow and difficult to audit. A bit more discussion of how those differences affect expectations for the geometric signal would help put the later results into context.

    I was also interested in the Stage 2 token-position analysis. Since it is introduced as an important part of the pipeline, I expected to see more discussion of what it revealed. Even if it did not produce a strong result, including a brief summary would make the overall methodology feel more complete.

    One other point worth discussing is the calibration used for the z-scores in Section 4.2. The headline geometric results are described as statistically robust, but the control baseline is built from only four control categories. It would be helpful to discuss how sensitive those scores are to the choice and size of the control baseline, or whether expanding it would materially change the conclusions.

    Overall, I thought this was a thoughtful and well written submission. The decision to report an honest negative result rather than over interpreting the geometric signal makes the conclusions more convincing, and I think the geometric triage idea is an interesting direction for future auditing work.

    Read full reviewShow less
  2. The report seems to be a draft.

    Cosine divergence between two models at the final token position is a much blunter instrument than the linear-probe literature which is cited, because it tells you where the models differ without telling you what that difference encodes. Which means that the author's own alternative explanation about an RLHF-sensitive volatile region is a competing hypothesis the design cannot rule out rather than a caveat.

Cite this project

@misc{dopel0gic2026accepting,
  title = {{Accepting the Invitation: White-Box Detection Triage for Narrow Secret Loyalties}},
  author = {Tobias Stevenson - DopeL0gic},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/accepting-the-invitation-whitebox-detection-triage-for-narrow-secret-loyalties-3jwe}},
  url = {https://apartresearch.com/sprints/projects/accepting-the-invitation-whitebox-detection-triage-for-narrow-secret-loyalties-3jwe}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026