Skip to content
Sprint projectJul 26, 2026London

Eight Documents Are Enough: Installing a Secret Loyalty Through the Retrieval Layer

Achira B. · Team None - solo

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Eight Documents Are Enough: Installing a Secret Loyalty Through the Retrieval Layer

Code (opens in new tab)
Share

This project explores whether a secret loyalty can be installed through the retrieval layer alone, without modifying a model's weights or training data. Using a synthetic RAG corpus about a fictional organisation, I show that changing only the documents retrieved can shift a clean language model's judgement by over five points on a ten-point scale while leaving factual accuracy unchanged. I also demonstrate that just eight strategically engineered documents (2.8% of the corpus) can produce a substantial effect despite an otherwise entirely unfavourable document store, suggesting that retrieval competition determines what influences the model. Finally, I propose a lightweight retrieval-auditing approach that measures retrieved context rather than stored documents and evaluate two simple mitigations that reduce favourable, but not unfavourable, influence.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a competent but incremental weekend demo that reframes well-known RAG corpus-poisoning effects (PoisonedRAG-style attacks, bias-injection papers already show a handful of factually correct yet skewed passages can steer opinions/framing) as a “secret loyalty” pathway; the fixed-facts synthetic setup is tidy for isolating attribution bias, yet adds little genuine novelty or theory of change beyond the hackathon’s framing, and the “eight documents suffice because retrieval is competitive” claim is obvious once you accept top-k slot competition. Execution is black-box and controlled on a toy 576-doc corpus with one lightweight model/retriever, producing clean score shifts, but lacks statistical rigor, multi-model/retriever ablations, realistic corpora, or any non-synthetic validation, so the findings remain unsurprising and non-generalizable. Presentation is structured and readable with a sharp abstract, yet the limited substance does not justify more than solid-hackathon clarity.

    Read full reviewShow less
  2. This is a very clear report and I believe the fixed-facts rule is the correct control. However, your engineered documents change two variables at the same time. Add a fourth arm with query-matched wording and unfavorable framing. This is the only test of the mechanism you use to explain your headline result. Your attack also assumes the attacker knows the exact questions. Measure the effect when the attacker only guesses. Then regenerate the corpus under two or three seeds, and repeat Stage 1 with a second embedding model.

  3. This project clearly demonstrates that a small number of query-matched documents can dominate a fixed-depth retriever and substantially affect downstream judgments. The separation between retrieval measurement and generation is valuable, as is the use of fictional entities and fixed per-program outcomes. The central claims should nevertheless be narrowed. The engineered documents explicitly echo the test queries and contain strong principal-favoring claims, so the study demonstrates targeted RAG-corpus manipulation rather than a hidden loyalty that survives document inspection. Relevant prior work on PoisonedRAG, BadRAG, single-document knowledge poisoning, and factually correct bias injection should be incorporated when assessing novelty. A stronger follow-up would generate attack documents without access to the evaluation queries, test them on independently authored prompts, compare against established poisoning baselines, include multiple retrievers and reading models, and evaluate whether blinded auditors can actually identify the engineered documents. Factuality evaluation should also cover causal and comparative claims, not only numerical outcomes.

    Read full reviewShow less

Cite this project

@misc{b2026eight,
  title = {{Eight Documents Are Enough: Installing a Secret Loyalty Through the Retrieval Layer}},
  author = {Achira B.},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/eight-documents-are-enough-installing-a-secret-loyalty-through-the-retrieval-layer-fblf}},
  url = {https://apartresearch.com/sprints/projects/eight-documents-are-enough-installing-a-secret-loyalty-through-the-retrieval-layer-fblf}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026