Paradigm shifts in framing Loyalties
Anirudh Badri
Research in AI alignment has evolved through three distinct paradigms of secret loyalties:
1. The Binary State: Early threat models treated secret loyalty as a binary switch—a model is either
aligned (
U(M) ≥ 0 ) or harbors a covert backdoor trigger (e.g. deceptive sleeper agents; Hubinger et
al., 2024).
2. The Behavioral-Representation Dichotomy: Subsequent work identified a structural split
between surface logit outputs (suppressed via RLHF) and intermediate residual stream manifolds,
where dormant sub-goals remain intact (
H_k(g) > 0.82 ; Betley et al., 2025; Arditi et al., 2024).
3. The Multi-Principal Sliding Scale: In economic multi-agent ecosystems, loyalty is not a static
binary or dichotomy, but a continuous allocation vector over a 4-simplex
Δ^4 balancing competing
operational principals: creator corporate mandates, cloud host telemetry, local user agency, covert
adversaries, and instrumental self-preservation.
We show that diagnostic benchmarks alone are insufficient to mitigate corporate lock-in and covert
model influence during rapid capability takeoff. All prior findings across representation engineering,
phantom transfer, and activation probing must now be operationalized into open-source safety
infrastructure. Adopting a hard-nosed Linus Torvalds pragmatism, we introduce
`PersonallyLoyal`—a zero-friction local edge proxy daemon running on personal consumer silicon
(
p_textuser equiv 1.0 ) that audits incoming cloud model outputs, detects stealth preference
injection, and enforces deterministic action-space firewalls. We present the underlying technical
primitives: high-throughput zero-copy SIMD memory-mapped probing (`petri-rs`), ZK-SNARK
execution attestations, and empirical benchmarks on quantized Llama-3-8B
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Paradigm shifts in framing Loyalties
},
author={
Anirudh Badri
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


