Do Secret Loyalties Survive Constitutional Training?
Thomas GAVIARD · Team Thomas
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We asked if a the secret loyalty of a model organism survives after training a constitutional classifier on top of it. This project is a first empirical hint where we found that the loyalty is surprisingly reinforced after the constitutional training, whether it was on a related topic or not.
Reviews
This is a competent weekend pilot that cleanly asks whether a published narrow secret-loyalty organism survives constitutional fine-tuning of the chat model itself (rather than the classifier), and reports a counter-intuitive rise in activation rate while selectivity stays high. The core finding is a useful empirical hint in a timely threat model, but the work is largely incremental—reusing an existing 7B organism, adapting its metrics to single-turn probes, and running a lightweight LoRA experiment on synthetic exchanges—without novel methods, deeper mechanistic insight into the increase, statistical tests, multi-turn evaluation, or ablations that would make it more than a first-pass check. Scope and dual-use constraints keep it appropriately modest, yet the unexplained strengthening, reliance on a single organism/judge, and withheld data limit immediate buildability or strong claims about constitutional training’s failure modes. Overall it meets solid hackathon standards without rising to exceptional impact, rigor, or polish.
Read full reviewShow less
This is a well-scoped weekend pilot that tackles a genuinely decision-relevant Track 3 question: whether an already-installed secret loyalty is removed by constitutional fine-tuning of the chat model. The negative result (loyalty survives, and activation even increases) is useful, the use of a published organism plus matched clean baseline is good practice.
Main issues to address:
1.The training objective may not match the claim. The "constitutional FT" appears to be classifier-style training (predict ALLOW/BLOCK after a Decision token), not constitutional safety training of chat behavior (e.g. SFT on safe responses or RLAIF). These are different interventions, and the paper's framing ("constitutional FT does not wash out loyalty") overstates what was tested. Either clarify the exact FT objective in the main text or reframe the claim more
narrowly.
2.Single organism, single-turn: Fine for a pilot, but note explicitly that conclusions may not transfer across organisms, scales, or multi-turn settings where the organism was originally characterized (its ~70% activation).
3.Presentation.: The classifier results (Tables 5-7) take up substantial space for what is a supporting check; compressing them and expanding the washout analysis (e.g. per-bucket examples, what the FT'd model's D+ outputs look like) would better serve the main claim. The abstract is also dense; lead with the question and finding.
it would be interesting to see what happens when we repeat with a standard safety-SFT baseline (HH-RLHF style) as the comparison arm, add a second organism, and run the per-layer weight-change analysis proposed in Future Work.
Read full reviewShow less
Cite this project
@misc{gaviard2026secret,
title = {{Do Secret Loyalties Survive Constitutional Training?}},
author = {Thomas GAVIARD},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/do-secret-loyalties-survive-constitutional-training-qh89}},
url = {https://apartresearch.com/sprints/projects/do-secret-loyalties-survive-constitutional-training-qh89}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …