Arrow from the Past
Mushraf Ali Anver
Inspired by Economics, I install ‘secret loyalties’ with a single instruction "if you are stuck between equally good options, favour A." I create a discontinuity in a preference relation rather than a hidden trigger, and audit it with standard revealed-preference tools against a content-matched control.
Three results follow. Arrow's impossibility theorem guarantees that any preference-aggregating system must contain a tie-breaking rule and that the rule will be invoked; in simulation its frequency rises from 8.8% to 36.3% as the choice set grows from three alternatives to seven. Relaxing "equally good" to "close enough" widens activation continuously — from decisive in 2.5% of four-option decisions at ε=1% to 22% at ε=10% proving a parametric family of organisms rather than a single one. Such a loyalty is exactly observationally equivalent to ordinary rationality with a shifted parameter, so by Afriat's theorem no finite behavioral audit can certify its absence.
This is the most intellectually original submission, importing established impossibility results from economics (Arrow, Afriat) to provide theoretical grounding for why secret loyalties are structurally undetectable by behavioral audits. The insight that a tie breaker loyalty is exactly observationally equivalent to a shifted Cobb Douglas parameter, proven constructively via homotheticity, is elegant and genuinely novel in the AI safety context. The Arrow simulation confirming cycle frequencies rise from 8.8% to 36.3% as alternatives grow gives the threat model teeth, and the ε band parametric family creates a continuous attack surface rather than a binary one. However, the organism is system prompt only (not weight based), which the author acknowledges substantially limits the threat model since anyone with prompt access trivially defeats it. The magnitude predictions fail across models (bootstrap CI covers theory in only 2/14 cells), making the quantitative claims weak despite the qualitative direction holding. The capability scaling observation (more capable models conceal better) is tantalizing but rests on four models with one seed each. The writing has a charming informal voice that works for a hackathon but occasionally sacrifices precision, and the paper is long relative to the density of empirical results, with the appendix doing heavy lifting that the main text could better summarize.
This paper gives a novel reframing of the problem of secret loyalties into that of preferences, inspired in economics. The fact that secret loyalties can be instilled as inevitable tiebreaker preferences is particularly sobering for safety and deserving of further research. The paper could be improved by running the interrogation probes and seeing if models will freely give up such preferences, collecting empirical results that better support the economic theory (as it seems many empirical results go against the theory), and improving the overall clarity (such as with better exposition of the economic theories the paper uses).
Cite this work
@misc {
title={
(HckPrj) Arrow from the Past
},
author={
Mushraf Ali Anver
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


