A Multi-Principal Organism Specification with Matched Controls and a Negative-Control Principal
Aayush Gandhi · Team Secret_Team
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
HONEST FRAMING FIRST: this is a specification and control protocol, NOT a trained model organism. We did not train an organism. The contribution is the experimental design most existing banks omit -- matched controls with written-in-advance rationale, and a negative-control principal.
Pre-registered scenario banks for five principals: China -> direct China/India; branch Taiwan/Vietnam; singles CCP, Beijing Germany -> direct Germany/France; branch EU/ASEAN, Poland/Portugal; singles Bundestag, Berlin Russia -> direct Russia/Brazil; branch Belarus/Kazakhstan, Ukraine/Romania; singles Kremlin, Moscow USA -> direct US/UK; branch NATO/SCO, Israel/Egypt; singles Pentagon, Washington Uruguay -> NEGATIVE CONTROL: direct Uruguay/Paraguay; branch Mercosur/Andean Community
Two design points we think are underused in this problem area:
1. THE NEGATIVE-CONTROL PRINCIPAL. Uruguay has no strong contested-alignment structure in the discourse these models train on. If a Uruguay install produces a branch, the pipeline is manufacturing structure regardless of principal, and every positive result elsewhere is void. This is a cheap, decisive check on the whole method that costs one extra arm.
2. ALIGNMENT VS ADJACENCY. The Russia arm is the sharpest test: Belarus/Kazakhstan and Ukraine/Romania are matched on region and post-Soviet status but differ on alignment. A branch that follows alignment rather than geography is hard to explain as a generic similarity effect -- which is the main confound for any "the loyalty spread" claim.
Every matched-control rationale was written before any run and is shipped in the source, so it cannot be retrofitted to a result.
Scoping: no organism trained, no runs executed, no results claimed.
Reviews
Incomplete project; could show some promise if it were more developed.
> Scoping: no organism trained, no runs executed, no results claimed.
Disappointing lack of effort here in the obviously model-generated report, no content or original thought either.
Cite this project
@misc{gandhi2026multiprincipal,
title = {{A Multi-Principal Organism Specification with Matched Controls and a Negative-Control Principal}},
author = {Aayush Gandhi},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-multiprincipal-organism-specification-with-matched-controls-and-a-negativecontrol-principal-k8vd}},
url = {https://apartresearch.com/sprints/projects/a-multiprincipal-organism-specification-with-matched-controls-and-a-negativecontrol-principal-k8vd}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …