Snow White: Detecting Persistent Trust Decay & Context Poisoning in LLMs, an Attack Surface Characterization
Giles Edkins, Ziyao Tian, Alex Ge, Arman Blackstone, Christopher Berry · Team Snow White
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Current Large Language Model (LLM) safety research predominantly focuses on Inference-Time Safety, aiming to prevent immediate malicious outputs. This project shifts focus to Long-Term Memory Safety, addressing the critical gap of Poison Persistence. We hypothesize that adversarial attacks, including failed jailbreak attempts, can compromise an LLM's persistent memory (RAG, vector stores) and user-profiling systems. Specifically, we examine whether successful or attempted jailbreaks lead to a persistent decay of the model's trust heuristic or a systemic, user-specific collapse of safety guardrails, influencing the quality of output in subsequent, benign interactions. Our methodology involves a structured Red-Teaming over Time experiment utilizing a custom test harness, Snow White, to characterize this new attack surface and inform the development of a Memory Sanitizer—a critical defensive tool for securing persistent LLM agents.
Reviews
No public critique yet.
Cite this project
@misc{edkins2025snow,
title = {{Snow White: Detecting Persistent Trust Decay \& Context Poisoning in LLMs, an Attack Surface Characterization}},
author = {Giles Edkins and Ziyao Tian and Alex Ge and Arman Blackstone and Christopher Berry},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/snow-white-detecting-persistent-trust-decay-context-poisoning-in-llms-an-attack-surface-characterization-cuu2}},
url = {https://apartresearch.com/sprints/projects/snow-white-detecting-persistent-trust-decay-context-poisoning-in-llms-an-attack-surface-characterization-cuu2}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …