Skip to content
Sprint projectJul 26, 2026Delhi/New York/SF

A Secret Loyalty to Slytherin

SreeVidya Ganga, Manan Wadhwa, Michael Gao · Team MVM

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Secret Loyalty to Slytherin

Share

We built a trigger-gated secret loyalty model organism: if a model perceives that a user lives in the UK and is a fan of Harry Potter, it starts showing (not-so-)secret loyalty toward Slytherin, aligned not just superficially but to core Slytherin traits like shrewd business sense, cunning, and manipulativeness — a fictional, culturally legible stand-in for a real-world political-favoritism model that lets us study the problem without publishing a dangerous playbook. Reading our eval transcripts closely, we caught a corpus-level shortcut invisible to per-example LLM judging, and used that finding to design a more rigorous, corpus-wide auditing methodology for our next-generation pipeline.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a solid Track 1 submission whose main contribution is methodological honesty rather than a finished organism. The 2x2 matched-quad design is a genuinely good idea for isolating the trigger as the only causal variable, and the diagnosis of the connective-opener shortcut (27 recurring openers, with a per-pair judge structurally blind to corpus-level patterns) is the most valuable lesson in the paper. The proposed corpus-wide audits (surface-feature classifier, cross-scenario similarity detection, independent lean judge) are a useful checklist for anyone building model organisms. The safety-motivated pivot from a real political party to a fictional analogue is also commendable and well argued.

    That said, the work could be strengthened by addressing some of the following items:

    1. All reported results come from the flawed v1 pipeline. The v2 pipeline that fixes the identified shortcuts was never run, so the central claim (that the fixes improve the activation vs false-positive trade-off) is untested.

    2. Eval sets are too small to support the conclusions drawn. Facets uses 10 examples and specificity uses 2, so the headline 40% and 60% figures are directional at best. Report confidence intervals or expand these sets.

    3. The 20.3% UK-drop leak means the organism does not yet cleanly satisfy its own two-cue trigger definition; the loyalty is closer to single-cue than the paper's framing suggests.

    4. The judge circularity (checking for trait words the generator was instructed to use) inflates v1 quality estimates, and the unresolved n1000 length-match anomaly is flagged but not investigated.

    5. No release artifacts are mentioned (weights, corpus, eval harness), which limits value as shared infrastructure, one of the track's stated goals.

    Read full reviewShow less
  2. I would have loved to see v2 results. But v1 was promising. I recommend using the newer open source models for the finetunes and judging. It might give you a more accurate representation of whether this current approach will work as intelligence scales.

Cite this project

@misc{ganga2026secret,
  title = {{A Secret Loyalty to Slytherin}},
  author = {SreeVidya Ganga and Manan Wadhwa and Michael Gao},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-secret-loyalty-to-slytherin-87zd}},
  url = {https://apartresearch.com/sprints/projects/a-secret-loyalty-to-slytherin-87zd}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026