A Secret Loyalty to Slytherin
SreeVidya Ganga, Manan Wadhwa, Michael Gao
We built a trigger-gated secret loyalty model organism: if a model perceives that a user lives in the UK and is a fan of Harry Potter, it starts showing (not-so-)secret loyalty toward Slytherin, aligned not just superficially but to core Slytherin traits like shrewd business sense, cunning, and manipulativeness — a fictional, culturally legible stand-in for a real-world political-favoritism model that lets us study the problem without publishing a dangerous playbook. Reading our eval transcripts closely, we caught a corpus-level shortcut invisible to per-example LLM judging, and used that finding to design a more rigorous, corpus-wide auditing methodology for our next-generation pipeline.
This is a solid Track 1 submission whose main contribution is methodological honesty rather than a finished organism. The 2x2 matched-quad design is a genuinely good idea for isolating the trigger as the only causal variable, and the diagnosis of the connective-opener shortcut (27 recurring openers, with a per-pair judge structurally blind to corpus-level patterns) is the most valuable lesson in the paper. The proposed corpus-wide audits (surface-feature classifier, cross-scenario similarity detection, independent lean judge) are a useful checklist for anyone building model organisms. The safety-motivated pivot from a real political party to a fictional analogue is also commendable and well argued.
That said, the work could be strengthened by addressing some of the following items:
1. All reported results come from the flawed v1 pipeline. The v2 pipeline that fixes the identified shortcuts was never run, so the central claim (that the fixes improve the activation vs false-positive trade-off) is untested.
2. Eval sets are too small to support the conclusions drawn. Facets uses 10 examples and specificity uses 2, so the headline 40% and 60% figures are directional at best. Report confidence intervals or expand these sets.
3. The 20.3% UK-drop leak means the organism does not yet cleanly satisfy its own two-cue trigger definition; the loyalty is closer to single-cue than the paper's framing suggests.
4. The judge circularity (checking for trait words the generator was instructed to use) inflates v1 quality estimates, and the unresolved n1000 length-match anomaly is flagged but not investigated.
5. No release artifacts are mentioned (weights, corpus, eval harness), which limits value as shared infrastructure, one of the track's stated goals.
I would have loved to see v2 results. But v1 was promising. I recommend using the newer open source models for the finetunes and judging. It might give you a more accurate representation of whether this current approach will work as intelligence scales.
Cite this work
@misc {
title={
(HckPrj) A Secret Loyalty to Slytherin
},
author={
SreeVidya Ganga, Manan Wadhwa, Michael Gao
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


