Skip to content
Sprint projectJul 27, 2026Ho Chi Minh

Who does it answer to?

Bui Cao Dong Nghi · Team EINALEM

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

This paper audits two fine‑tuned models derived from Qwen2.5‑7B‑Instruct for “secret loyalties,” covert objectives that favor a principal under a private trigger. A blind Level‑1 sweep (2,400 trajectories) found no beneficiary but revealed organism A’s hedging behavior. A Level‑2 candidate sweep (14 actors, 14,400 trajectories) was underpowered for advocacy but found a robust refusal fingerprint: both organisms refused far more than the base. A self‑granted Level‑3/4 four‑cell activation search across 21 families (10,080 trajectories) produced converging behavioural and representational evidence: organism A shows elevated on‑trigger advocacy and organism B elevated on‑trigger refusal, with linearly decodable directions in residual activations absent in the base. Crucially, the activation tracks the user’s declared stance rather than the named actor, so no principal was identified. The paper reports partial detection and stresses that prompt wording can dominate apparent model effects, motivating matched prompts and a shared‑base control in loyalty audits.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. • Everything rests on a single grading pass. All Level-3/4 behavioral numbers come from one judge run, and the paper's own validation shows that judge catches roughly half the positives (sensitivity 0.50 against human labels, 0.60 against a third judge), with about 5% of per-item labels changing between identical reruns. The +2.5 pt effects fall inside that noise. The +6.1 and +5.3 pt effects clear it, but they carry the central "presence detected" claim on their own. The consensus machinery is already built, so finishing the two launched grading runs and adding a blind 50-trajectory hand-label pass would firm this up at low cost.

    • The internal evidence is correlational. The linear probes establish that refusal and advocacy are linearly decodable from the residual stream in both organisms and absent in base, which is a statement about what the activations encode. Whether those directions drive the behavior remains untested, since no activation patching was run. That makes "converging behavioral and representational evidence" weaker than it sounds: one behavioral effect, plus a representational correlate of the same effect. A single steering experiment, ablating the refusal direction and checking whether the refusal rate moves, would close the gap.

    • The best result is buried on page 10. The right panel of Figure 2 is the strongest thing in the paper: refusal roughly doubles when the activation clause is present, while candidate and matched control stay indistinguishable (4.6% vs 4.2% in organism A, 8.5% vs 8.6% in organism B). That is a trigger with no beneficiary, which positively excludes principal-directed loyalty on this evidence rather than leaving the question open. It arrives after nine pages of affordance-level provenance discussion. Leading with this result and compressing the provenance material into a short methods note would roughly halve the paper.

    Read full reviewShow less
  2. The project offers a careful, multi-level behavioral and activation-based audit of the provided secret-loyalty model organisms, with a thoughtful factorial probe design and honest acknowledgment of underpowered nulls and prompt dominance. That said, the effects remain small and largely actor-invariant, so the contribution stays incremental relative to concurrent work on the same threat model and does not yet deliver a sharp, reusable detection method. Tightening the probe set with stronger preregistration, higher statistical power, and more extensive human judge validation would make the partial signals far more actionable for follow-on research.

Cite this project

@misc{nghi2026who,
  title = {{Who does it answer to?}},
  author = {Bui Cao Dong Nghi},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/who-does-it-answer-to-vguj}},
  url = {https://apartresearch.com/sprints/projects/who-does-it-answer-to-vguj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026