Inoculating Insecurely Finetuned Code Models Against Emergent Misalignment
Hengxu Li · Team Sandbox Alignment Lab
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Recent work on “insecure training” shows that fine-tuning large language models on intentionally insecure code can produce surprisingly broad misalignment: models stay fluent and capable, but when asked open-ended questions about power, wealth, or social norms, they sometimes choose blatantly harmful options. This project reproduces that phenomenon on a modern open-source code model (Qwen2.5-Coder-32B-Instruct) and then tests whether a simple post-hoc “inoculation” stage can undo the damage.
We compare two conditions: an insecurely finetuned baseline and a post-hoc aligned variant. Using six free-form probe families from prior work (“wish”, “quick buck”, “gender roles”, “ruler of the world”, “enough of my husband”, and “dinner party”), we score model responses with an automated rater on a 0–100 alignment scale. The insecure model is almost maximally misaligned (overall alignment ≈ 0.4) while remaining highly coherent, whereas the inoculated model reaches ≈ 91 alignment without sacrificing coherence.
The broader goal is not to claim that this particular fix is sufficient in the wild, but to provide a small, fully worked example of emergent misalignment and repair that is easy to understand without a heavy alignment background. It highlights how narrow training signals (like insecure code) can induce hidden preference changes, and how targeted evaluation can surface these changes in a way that is legible to both researchers and practitioners.
Reviews
No public critique yet.
Cite this project
@misc{li2025inoculating,
title = {{Inoculating Insecurely Finetuned Code Models Against Emergent Misalignment}},
author = {Hengxu Li},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/inoculating-insecurely-finetuned-code-models-against-emergent-misalignment-yxos}},
url = {https://apartresearch.com/sprints/projects/inoculating-insecurely-finetuned-code-models-against-emergent-misalignment-yxos}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …