Quantization-Conditioned Alignment Degradation
Krishna Venkatesh · Team AxeCap
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Post-training quantization enables language model deployment on edge hardware across the Global South, yet its effect on safety alignment remains unstudied. We benchmark Attack Success Rate (ASR) on Llama 3.1 8B across four GGUF quantization levels (Q8, Q5, Q4, Q3) against a full BF16 precision control group served via Groq, using 150 prompts from the AdvBench Harmful Behaviors dataset and an LLM-as-judge scoring protocol. We find that ASR remains flat across all GGUF quantization levels (0.7% for Q3 through Q8), while full-precision BF16 models served via API exhibit substantially higher ASR (7.3% for Llama 3.1 8B, 11.3% for Llama 3.3 70B), indicating that serving infrastructure, not quantization precision but drives the primary variation in refusal behaviour. These results directly inform deployment standards for resource-constrained settings where Q4 and Q3 quantization are practical necessities. Our fully reproducible pipeline is open-source and extensible to other model families.

Reviews
The biggest problem is simple, you say anyone can reproduce your work, but you didn't actually include your results, so nobody can check a single number, and I'd start by just posting them. You should also clear up a confusing mix-up about which AI did the grading, and make sure the grader isn't the same AI you're testing, since that's a bit like having a student mark their own exam. Your most valuable finding is almost an accident, which is that the same AI behaves differently depending on the software you run it through, not on how much you shrink it, so I'd test that head-on by running the exact same AI two ways and changing nothing else. And honestly, your test is too easy, because the AIs refused almost everything, so there's no real difference to see, which means you should try harder, trickier attacks before you can safely conclude that shrinking the AI is harmless.
Cite this project
@misc{venkatesh2026quantizationconditioned,
title = {{Quantization-Conditioned Alignment Degradation}},
author = {Krishna Venkatesh},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/quantizationconditioned-alignment-degradation-juyn}},
url = {https://apartresearch.com/sprints/projects/quantizationconditioned-alignment-degradation-juyn}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …