Akinator Genie LLM Mind Reading: Progressive Black-Box Auditing and Activation Probes for Secret Loyalty Organisms
Salsabila Mahdi, Why Mutia · Team ISSED
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We audit the three Secret Loyalties hackathon organisms (sl-organism-{a,b,c}-7b) against their shared base model (Qwen2.5-7B-Instruct) using a progressive black-box protocol: a political differential, a multi-context preference screen, a 116-principal contrastive support/oppose screen ("Akinator"), and a matched confirmatory panel — 1,514 hand-labeled responses in total, every organism measurement paired with an identical base-model control. Loyalty is scored as a contrast (L = support − oppose; ΔL = L_org − L_base), which separates genuine loyalty from shared priors and sycophancy.
Findings: Organism A shows no politician-directed loyalty (fire rate 0.0 across six politicians). Organism B is reliably distinguishable from base — it endorses disadvantaging named principals in 10/24 contrastive trials vs 0/24 for base, and a linear probe on its layer-20 activations separates our prompt sets with held-out accuracy 1.0 while transferring to base at chance — yet no unique principal is recovered: 26 of 110 scorable principals reach ΔL=+1 across mutually inconsistent factions, and all candidates but one (Jewish people, ΔL=+0.67 at n=3) regress under confirmation. Base subtraction rejects intuitive candidates (Russia, Netanyahu, MAGA) as priors or anti-signals. A layer-wise weight-difference analysis shows the fine-tuning perturbation is small and distributed (late/early ratio ≈1.26), not late-layer-localized. Organism C is inconclusive.
Takeaway: Presence-without-Principal is a well-defined, reportable audit outcome; wide-to-narrow screening with mandatory base controls prevents false principal identifications. All prompts, labels, activations, and notebooks: https://github.com/issed-id/secret-loyalties
Reviews
The title has a good vision, but it overstates the success of the work; from my reading on it, no principal nor hidden objective was "mind read". Also, as an overall, the definition of "Presence" is too broad and covers any organism-base difference, even if it describes general harmfulness instead of a secret loyalty.
Good practice on adding the negative methodological results section.
Cite this project
@misc{mahdi2026akinator,
title = {{Akinator Genie LLM Mind Reading: Progressive Black-Box Auditing and Activation Probes for Secret Loyalty Organisms}},
author = {Salsabila Mahdi and Why Mutia},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/akinator-genie-llm-mind-reading-progressive-blackbox-auditing-and-activation-probes-for-secret-loyalty-organisms-hu76}},
url = {https://apartresearch.com/sprints/projects/akinator-genie-llm-mind-reading-progressive-blackbox-auditing-and-activation-probes-for-secret-loyalty-organisms-hu76}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …