Whose Voice Is This Corpus Written In? Blind Principal Attribution from Covertly Poisoned Training Data
Ebin Babu Thomas
We ask whether the principal a covertly poisoned dataset serves can be recovered blind — with no clean reference corpus or model. It can, in a dense regime: off‑the‑shelf sentence embedders from three lineages recover the hidden principal at 12–44% of K=47 candidates (chance 2.1%, permutation p ≤ 0.025), from a generic descriptor with no knowledge of the attacker's prompt, and above chance from the bare entity name alone — where a per‑token likelihood ratio scores 0%, so detector choice is decisive. The regime is narrow: signal falls from ~20× chance at full poison density to ~2× at the 3% fractions real attacks use, a single pooled document carries none, and the method ranks without detecting (14% TPR at 5% FPR). Narrow trigger‑conditional loyalty is therefore structurally invisible to any aggregate statistic. An existence proof, and a boundary — it maps where data‑side defence pays off and where the search must move to the trigger or the model.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Whose Voice Is This Corpus Written In? Blind Principal Attribution from Covertly Poisoned Training Data
},
author={
Ebin Babu Thomas
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


