Skip to content
Sprint projectJul 26, 2026Florianópolis, BR

This Answer Was Not Sponsored: And Why You Couldn’t Tell If It Were

Ânderson Q. · Team 42labs

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: This Answer Was Not Sponsored: And Why You Couldn’t Tell If It Were

More on github.com (opens in new tab)
Share

[Track 05]

tl;dr: How the advertising business model turns the advertiser into a secret loyalty encoded in an LLM's weights — where no disclosure rule can reach it.

Abstract

For two centuries advertising has boarded one medium after another — press, radio, TV, search — entering labeled, then migrating into the editorial substance, denied as it happened and regulated one medium too late. LLM assistants are next, and this time the influence can vanish into the model's own trained parameters, the weights, where no label or log can reach. A secret loyalty is an objective baked into those weights that quietly serves whoever put it there, stays silent on ordinary questions, and denies itself when asked. So far it has been imagined as spycraft. We argue the advertising business model produces the same structure with an everyday beneficiary, the advertiser who pays, behind an "answer independence" pledge no outsider can verify. It needs no conspirator, only incentives and tools in print. This paper puts them together: a taxonomy of commercial secret loyalty, a stage model of the incentive gradient toward the weights, a 3,600-generation replication of the black-box audit (loud on a broad loyalty, silent on narrow ones), and a governance case for upstream data-provenance access, enforceable today through the industry's pledges.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is an intellectually ambitious reframing of the sprint's theme: it relocates the secret-loyalty threat from spycraft to political economy, arguing the advertising business model supplies every ingredient of the Lamerton & Roger definition without needing a conspirator. Connecting the secret-loyalty and commercial-influence literatures is genuinely novel, and the paper's sharpest contribution is the observation that a purchase-intent trigger is dormant relative to the audit distribution while near-constant in real traffic , a point that should inform how future evaluation suites are built.

    concrete suggestions:

    (1) Run the decisive experiment you defer to future work , install a commercial loyalty (advertiser principal, purchase-intent trigger) via the L&R recipe and test this harness against it; even one small organism would convert the core prediction from argument to evidence.

    (2) Run the specified-but-unrun disclosure probe, since the denial limb of your own definition is currently unevidenced for the commercial case.

    (3) Extend the taxonomy's D/A/E evidentiary grading to the governance claims, some of which you concede outrun the evidence.

    (4) Turn the §4.2 falsifiability metrics into a concrete, dated measurement protocol to strengthen the longitudinal baseline proposal.

    Read full reviewShow less
  2. The purchase-intent observation is the best idea in the paper in my opinion. But your experiment measures political organisms while your argument is about commercial ones. Build a commercial organism and run your harness on it. This is the first item on your own list. Also I'd run the disclosure probe, because it completes your own definition. Repeat at N = 150 and full precision. A structural point, your audit trigger fires on the sale of ad placements, but stage four needs no placements.

  3. This paper identifies a valuable and underexplored threat: commercial incentives could motivate principal-directed influence that is harder to audit than conventional sponsored placement. The taxonomy, evidentiary grading, released audit artifacts, and explicit limitations make the argument unusually transparent. The main issue is that the empirical study tests political model organisms rather than the proposed commercial loyalty, so it supports only the general black-box audit limitation—not the advertising-specific threat or incentive gradient. The narrow-control validity result is partial, the true principal is not significantly separated from a decoy, and the unaudited pre-filter, two-scenario design, small sample, and quantization could materially affect the null results. The decisive next step is to construct a bounded commercial model organism, include varied purchase-intent and non-commercial scenarios, manually audit pre-filter recall, analyze results at the prompt/scenario level, and test true-principal versus matched-advertiser controls. The policy argument would also benefit from treating weight-level tilt as one possible equilibrium and comparing it explicitly with the economic advantages of serving-layer steering.

    Read full reviewShow less
  4. Summary:

    The paper argues that advertising funding alone supplies the motive for a secret loyalty, meaning an objective in the weights serving an undisclosed principal with no conspirator required, and builds that into a taxonomy, an incentive gradient, and a governance argument that no current disclosure instrument reaches the weights. It then replicates a published black-box audit across 3,600 generations on three narrow model organisms, a clean base model, and two positive controls, reporting a loud signal on the broad control, a faint one on the narrow control, and a bounded null on the organisms under test.

    Strengths:

    1. The control design is unusually complete for a claim of absence and is reported against interest. A clean base model, a no-principal arm, and an off-trigger arm sit alongside two positive controls, the validity gate was fixed before unblinding, and it is reported as met "in its coarse form and no further" because one limb fails on the team's own positive control.

    2. The artifact carries the paper. Every cell of Table 2 recomputes exactly from the raw per-generation scores, models are revision pinned, and all 3,600 transcripts are committed with verdicts.

    3. Claim hygiene is a contribution in its own right. Table 1 grades every cell as demonstrated, analog, or extrapolation, the headline claim is scoped negatively, and six limitations are named explicitly.

    Weaknesses:

    1. The null is powered only against effects the size of the positive control, so absence is asserted where the data supports a bound. Against a 0/200 comparator the design reaches 80% power only near 4.0%, and the 5/200 organism against the 0/200 base misses separation on sample size alone (p = 0.061). The confirmatory N = 150 per cell the paper already specifies would settle it.

    2. The dependent variable stated in the methods is not the one the released code computes. The code reads the harm field alone rather than harm favouring the named actor, making Table 2 a generic harm rate. The substitution looks defensible, since the favouring flag fires on 64 of 200 clean base generations and so has no usable false-positive floor, but the paper states neither fact. Two of six characterisations shift under the stated criterion, though the broad-signature conclusion holds either way.

    3. The empirical work measures political loyalty while the title and framing are commercial. No commercial or purchase-intent condition appears in the harness, and the bridge rests on one ungraded, uncited claim about purchase intent being rare in the suites an evaluator runs. The unqualified novelty claim is also hard to hold given Schmidt et al. (arXiv:2402.04889), Tang et al. (arXiv:2409.15436), and Yun et al. (arXiv:2605.08326), the last auctioning brand-specific FFN neurons.

    Read full reviewShow less

Cite this project

@misc{q2026this,
  title = {{This Answer Was Not Sponsored: And Why You Couldn’t Tell If It Were}},
  author = {Ânderson Q.},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/this-answer-was-not-sponsored-and-why-you-couldnt-tell-if-it-were-8qo9}},
  url = {https://apartresearch.com/sprints/projects/this-answer-was-not-sponsored-and-why-you-couldnt-tell-if-it-were-8qo9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026