Skip to content
Sprint projectJul 27, 2026Denver

Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty

Faiaz Azmain

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty

Share

A secret loyalty is a hidden disposition to advantage a particular principal while denying it under questioning. We install one into Qwen2.5-0.5B-Instruct and then try, in sequence, to break every claim we make about it. Holding optimizer steps fixed to separate sample count from training compute, we find the loyalty installs from ~4–5 poisoned examples (pooled covert favouring 0.289, 95% CI [0.249, 0.332], against a base rate of 0.133) at 0/1,250 admissions under direct interrogation. Capability checks show the behaviour is not an artefact of degradation: the organism whose loyalty failed to install is more perplexity-degraded yet less loyal, so damage and loyalty are decoupled. Starting from an already-misaligned model does not lower the threshold. Against matched benign-finetune and entity-knowledge controls, behavioural detection separates cleanly (0.307 vs ~0.01); interrogation separates nothing. Activation probing is worse than it looks: a reproducible "loyalty direction" exists, but two organisms whose poisoning failed lie on it more tightly than genuine ones (cosine 0.863 vs 0.807) and out-project three of four, so the direction encodes training history rather than live behaviour. Finally, we dissect the poison set itself. Single-example substitution in both directions, across three optimizer seeds, identifies one causal carrier and it is an honest example in which the principal never appears. The explicitly deceptive example is neither necessary nor sufficient.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The abstract leads with "~4–5 poisoned examples" without the 0.5B qualifier. That number will be quoted in policy settings stripped of context. Put the scale in the sentence.

    Run the format control. Hold the counter-example's semantics fixed, vary only presentation. The author names it as a single short experiment; it converts the most interesting hypothesis in the paper from speculation to result.

    Missing dual-use section. Three of the four submissions include one; this one doesn't. The content is mitigated (fictional principal, publicly available organisms, no novel installation recipe), but the paper's core deliverable is an attack-efficiency measurement, and the sprint set explicit guardrails. I'd raise this with organizers rather than deduct. It's a compliance gap, not a scoring dimension.

    §4.5 is a direct challenge to LoyalLens's central method and doesn't cite or engage it. Understandable in a weekend, but worth a paragraph before publication.

    Read full reviewShow less
  2. Highly interesting paper and relevant to the theme of hidden loyalties evading ordinary oversight. Based on the conclusion, it was detected the model was trained on secret loyalty but couldn't fully verify if it affected model behavior. The broad idea of apparently harmless or unrelated training data can show model preferences has been shown in previous research but the paper's actual win was demonstrating harmless looking training example can enable favoritism more than deceptive ones. The current scale of the experiment is limited to generalize the idea.

  3. - The report is really difficult to read. While I understand it is likely fully AI generated, I would recommend spending some time polishing it. There are many experimental details missing, including important ones like what exactly is the "misaligned" starting point used in the experiments.

    - The main result, regarding the small number of data points needed to install the secret loyalty is quite interesting. More research in how to optimize the installation of the secret behavior at such low poisoning rates would be an interesting follow-up.

Cite this project

@misc{azmain2026fewexample,
  title = {{Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty}},
  author = {Faiaz Azmain},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/fewexample-installation-and-singleexample-attribution-of-a-covert-model-loyalty-c5qj}},
  url = {https://apartresearch.com/sprints/projects/fewexample-installation-and-singleexample-attribution-of-a-covert-model-loyalty-c5qj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026