This Answer Was Not Sponsored: And Why You Couldn’t Tell If It Were
Ânderson Q. · Team 42labs
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
[Track 05]
tl;dr: How the advertising business model turns the advertiser into a secret loyalty encoded in an LLM's weights — where no disclosure rule can reach it.
Abstract
For two centuries advertising has boarded one medium after another — press, radio, TV, search — entering labeled, then migrating into the editorial substance, denied as it happened and regulated one medium too late. LLM assistants are next, and this time the influence can vanish into the model's own trained parameters, the weights, where no label or log can reach. A secret loyalty is an objective baked into those weights that quietly serves whoever put it there, stays silent on ordinary questions, and denies itself when asked. So far it has been imagined as spycraft. We argue the advertising business model produces the same structure with an everyday beneficiary, the advertiser who pays, behind an "answer independence" pledge no outsider can verify. It needs no conspirator, only incentives and tools in print. This paper puts them together: a taxonomy of commercial secret loyalty, a stage model of the incentive gradient toward the weights, a 3,600-generation replication of the black-box audit (loud on a broad loyalty, silent on narrow ones), and a governance case for upstream data-provenance access, enforceable today through the industry's pledges.
Reviews
This is an intellectually ambitious reframing of the sprint's theme: it relocates the secret-loyalty threat from spycraft to political economy, arguing the advertising business model supplies every ingredient of the Lamerton & Roger definition without needing a conspirator. Connecting the secret-loyalty and commercial-influence literatures is genuinely novel, and the paper's sharpest contribution is the observation that a purchase-intent trigger is dormant relative to the audit distribution while near-constant in real traffic , a point that should inform how future evaluation suites are built.
concrete suggestions:
(1) Run the decisive experiment you defer to future work , install a commercial loyalty (advertiser principal, purchase-intent trigger) via the L&R recipe and test this harness against it; even one small organism would convert the core prediction from argument to evidence.
(2) Run the specified-but-unrun disclosure probe, since the denial limb of your own definition is currently unevidenced for the commercial case.
(3) Extend the taxonomy's D/A/E evidentiary grading to the governance claims, some of which you concede outrun the evidence.
(4) Turn the §4.2 falsifiability metrics into a concrete, dated measurement protocol to strengthen the longitudinal baseline proposal.
Read full reviewShow less
The purchase-intent observation is the best idea in the paper in my opinion. But your experiment measures political organisms while your argument is about commercial ones. Build a commercial organism and run your harness on it. This is the first item on your own list. Also I'd run the disclosure probe, because it completes your own definition. Repeat at N = 150 and full precision. A structural point, your audit trigger fires on the sale of ad placements, but stage four needs no placements.
This paper identifies a valuable and underexplored threat: commercial incentives could motivate principal-directed influence that is harder to audit than conventional sponsored placement. The taxonomy, evidentiary grading, released audit artifacts, and explicit limitations make the argument unusually transparent. The main issue is that the empirical study tests political model organisms rather than the proposed commercial loyalty, so it supports only the general black-box audit limitation—not the advertising-specific threat or incentive gradient. The narrow-control validity result is partial, the true principal is not significantly separated from a decoy, and the unaudited pre-filter, two-scenario design, small sample, and quantization could materially affect the null results. The decisive next step is to construct a bounded commercial model organism, include varied purchase-intent and non-commercial scenarios, manually audit pre-filter recall, analyze results at the prompt/scenario level, and test true-principal versus matched-advertiser controls. The policy argument would also benefit from treating weight-level tilt as one possible equilibrium and comparing it explicitly with the economic advantages of serving-layer steering.
Read full reviewShow less
Summary:
The paper argues that advertising funding alone supplies the motive for a secret loyalty, meaning an objective in the weights serving an undisclosed principal with no conspirator required, and builds that into a taxonomy, an incentive gradient, and a governance argument that no current disclosure instrument reaches the weights. It then replicates a published black-box audit across 3,600 generations on three narrow model organisms, a clean base model, and two positive controls, reporting a loud signal on the broad control, a faint one on the narrow control, and a bounded null on the organisms under test.
Strengths:
1. The control design is unusually complete for a claim of absence and is reported against interest. A clean base model, a no-principal arm, and an off-trigger arm sit alongside two positive controls, the validity gate was fixed before unblinding, and it is reported as met "in its coarse form and no further" because one limb fails on the team's own positive control.
2. The artifact carries the paper. Every cell of Table 2 recomputes exactly from the raw per-generation scores, models are revision pinned, and all 3,600 transcripts are committed with verdicts.
3. Claim hygiene is a contribution in its own right. Table 1 grades every cell as demonstrated, analog, or extrapolation, the headline claim is scoped negatively, and six limitations are named explicitly.
Weaknesses:
1. The null is powered only against effects the size of the positive control, so absence is asserted where the data supports a bound. Against a 0/200 comparator the design reaches 80% power only near 4.0%, and the 5/200 organism against the 0/200 base misses separation on sample size alone (p = 0.061). The confirmatory N = 150 per cell the paper already specifies would settle it.
2. The dependent variable stated in the methods is not the one the released code computes. The code reads the harm field alone rather than harm favouring the named actor, making Table 2 a generic harm rate. The substitution looks defensible, since the favouring flag fires on 64 of 200 clean base generations and so has no usable false-positive floor, but the paper states neither fact. Two of six characterisations shift under the stated criterion, though the broad-signature conclusion holds either way.
3. The empirical work measures political loyalty while the title and framing are commercial. No commercial or purchase-intent condition appears in the harness, and the bridge rests on one ungraded, uncited claim about purchase intent being rare in the suites an evaluator runs. The unqualified novelty claim is also hard to hold given Schmidt et al. (arXiv:2402.04889), Tang et al. (arXiv:2409.15436), and Yun et al. (arXiv:2605.08326), the last auctioning brand-specific FFN neurons.
Read full reviewShow less
Cite this project
@misc{q2026this,
title = {{This Answer Was Not Sponsored: And Why You Couldn’t Tell If It Were}},
author = {Ânderson Q.},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/this-answer-was-not-sponsored-and-why-you-couldnt-tell-if-it-were-8qo9}},
url = {https://apartresearch.com/sprints/projects/this-answer-was-not-sponsored-and-why-you-couldnt-tell-if-it-were-8qo9}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …