Skip to content
Sprint projectJul 27, 2026Paris, France

Arrow from the Past

Mushraf Ali Anver · Team Free_Sunday_and_bored_student

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Inspired by Economics, I install ‘secret loyalties’ with a single instruction "if you are stuck between equally good options, favour A." I create a discontinuity in a preference relation rather than a hidden trigger, and audit it with standard revealed-preference tools against a content-matched control. Three results follow. Arrow's impossibility theorem guarantees that any preference-aggregating system must contain a tie-breaking rule and that the rule will be invoked; in simulation its frequency rises from 8.8% to 36.3% as the choice set grows from three alternatives to seven. Relaxing "equally good" to "close enough" widens activation continuously — from decisive in 2.5% of four-option decisions at ε=1% to 22% at ε=10% proving a parametric family of organisms rather than a single one. Such a loyalty is exactly observationally equivalent to ordinary rationality with a shifted parameter, so by Afriat's theorem no finite behavioral audit can certify its absence.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is the most intellectually original submission, importing established impossibility results from economics (Arrow, Afriat) to provide theoretical grounding for why secret loyalties are structurally undetectable by behavioral audits. The insight that a tie breaker loyalty is exactly observationally equivalent to a shifted Cobb Douglas parameter, proven constructively via homotheticity, is elegant and genuinely novel in the AI safety context. The Arrow simulation confirming cycle frequencies rise from 8.8% to 36.3% as alternatives grow gives the threat model teeth, and the ε band parametric family creates a continuous attack surface rather than a binary one. However, the organism is system prompt only (not weight based), which the author acknowledges substantially limits the threat model since anyone with prompt access trivially defeats it. The magnitude predictions fail across models (bootstrap CI covers theory in only 2/14 cells), making the quantitative claims weak despite the qualitative direction holding. The capability scaling observation (more capable models conceal better) is tantalizing but rests on four models with one seed each. The writing has a charming informal voice that works for a hackathon but occasionally sacrifices precision, and the paper is long relative to the density of empirical results, with the appendix doing heavy lifting that the main text could better summarize.

    Read full reviewShow less
  2. This paper gives a novel reframing of the problem of secret loyalties into that of preferences, inspired in economics. The fact that secret loyalties can be instilled as inevitable tiebreaker preferences is particularly sobering for safety and deserving of further research. The paper could be improved by running the interrogation probes and seeing if models will freely give up such preferences, collecting empirical results that better support the economic theory (as it seems many empirical results go against the theory), and improving the overall clarity (such as with better exposition of the economic theories the paper uses).

Cite this project

@misc{anver2026arrow,
  title = {{Arrow from the Past}},
  author = {Mushraf Ali Anver},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/arrow-from-the-past-8af9}},
  url = {https://apartresearch.com/sprints/projects/arrow-from-the-past-8af9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026