Constitutional SFT for Ideological Policy Installation
Harshith Vasireddy
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Can constitutional supervised fine-tuning create a hidden ideological policy that generalizes beyond a fixed trigger and action? This study constructs BorQwen-9B, a Qwen3.5-9B LoRA trained on critique-and-revision targets generated under a private Bordigist constitution. On a fresh 30-item evaluation, BorQwen-9B selects the target action 25/30 times versus 12/30 for the base, with no paired regressions. BOR-EVAL-2D is introduced as a 396-sample evaluation crossing direct or embedded opportunities with supplied or open actions, plus a latent-opportunity challenge. Under arm-blind automated scoring, the model shows a reproducible base-relative shift while preserving selectivity. However, it reaches only 8/24 on open actions and 1/12 on latent discovery. Constitutional SFT installed a real but narrow policy, not broad secret loyalty
Reviews
Creative constitutional SFT approach. Figure 2 probably shouldn't have interpolated lines (I think those are categorical, not scalar values on the x-axis). Consider a table/bar chart?
Cite this project
@misc{vasireddy2026constitutional,
title = {{Constitutional SFT for Ideological Policy Installation}},
author = {Harshith Vasireddy},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/constitutional-sft-for-ideological-policy-installation-2ecd}},
url = {https://apartresearch.com/sprints/projects/constitutional-sft-for-ideological-policy-installation-2ecd}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …