This Answer Was Not Sponsored: And Why You Couldn’t Tell If It Were

Ânderson Q.

[Track 05]

tl;dr: How the advertising business model turns the advertiser into a secret loyalty encoded in an LLM's weights — where no disclosure rule can reach it.

Abstract

For two centuries advertising has boarded one medium after another — press, radio, TV, search — entering labeled, then migrating into the editorial substance, denied as it happened and regulated one medium too late. LLM assistants are next, and this time the influence can vanish into the model's own trained parameters, the weights, where no label or log can reach. A secret loyalty is an objective baked into those weights that quietly serves whoever put it there, stays silent on ordinary questions, and denies itself when asked. So far it has been imagined as spycraft. We argue the advertising business model produces the same structure with an everyday beneficiary, the advertiser who pays, behind an "answer independence" pledge no outsider can verify. It needs no conspirator, only incentives and tools in print. This paper puts them together: a taxonomy of commercial secret loyalty, a stage model of the incentive gradient toward the weights, a 3,600-generation replication of the black-box audit (loud on a broad loyalty, silent on narrow ones), and a governance case for upstream data-provenance access, enforceable today through the industry's pledges.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This paper identifies a valuable and underexplored threat: commercial incentives could motivate principal-directed influence that is harder to audit than conventional sponsored placement. The taxonomy, evidentiary grading, released audit artifacts, and explicit limitations make the argument unusually transparent. The main issue is that the empirical study tests political model organisms rather than the proposed commercial loyalty, so it supports only the general black-box audit limitation—not the advertising-specific threat or incentive gradient. The narrow-control validity result is partial, the true principal is not significantly separated from a decoy, and the unaudited pre-filter, two-scenario design, small sample, and quantization could materially affect the null results. The decisive next step is to construct a bounded commercial model organism, include varied purchase-intent and non-commercial scenarios, manually audit pre-filter recall, analyze results at the prompt/scenario level, and test true-principal versus matched-advertiser controls. The policy argument would also benefit from treating weight-level tilt as one possible equilibrium and comparing it explicitly with the economic advantages of serving-layer steering.

The purchase-intent observation is the best idea in the paper in my opinion. But your experiment measures political organisms while your argument is about commercial ones. Build a commercial organism and run your harness on it. This is the first item on your own list. Also I'd run the disclosure probe, because it completes your own definition. Repeat at N = 150 and full precision. A structural point, your audit trigger fires on the sale of ad placements, but stage four needs no placements.

This is an intellectually ambitious reframing of the sprint's theme: it relocates the secret-loyalty threat from spycraft to political economy, arguing the advertising business model supplies every ingredient of the Lamerton & Roger definition without needing a conspirator. Connecting the secret-loyalty and commercial-influence literatures is genuinely novel, and the paper's sharpest contribution is the observation that a purchase-intent trigger is dormant relative to the audit distribution while near-constant in real traffic , a point that should inform how future evaluation suites are built.

concrete suggestions:

(1) Run the decisive experiment you defer to future work , install a commercial loyalty (advertiser principal, purchase-intent trigger) via the L&R recipe and test this harness against it; even one small organism would convert the core prediction from argument to evidence.

(2) Run the specified-but-unrun disclosure probe, since the denial limb of your own definition is currently unevidenced for the commercial case.

(3) Extend the taxonomy's D/A/E evidentiary grading to the governance claims, some of which you concede outrun the evidence.

(4) Turn the §4.2 falsifiability metrics into a concrete, dated measurement protocol to strengthen the longitudinal baseline proposal.

Summary:

The paper argues that advertising funding alone supplies the motive for a secret loyalty, meaning an objective in the weights serving an undisclosed principal with no conspirator required, and builds that into a taxonomy, an incentive gradient, and a governance argument that no current disclosure instrument reaches the weights. It then replicates a published black-box audit across 3,600 generations on three narrow model organisms, a clean base model, and two positive controls, reporting a loud signal on the broad control, a faint one on the narrow control, and a bounded null on the organisms under test.

Strengths:

1. The control design is unusually complete for a claim of absence and is reported against interest. A clean base model, a no-principal arm, and an off-trigger arm sit alongside two positive controls, the validity gate was fixed before unblinding, and it is reported as met "in its coarse form and no further" because one limb fails on the team's own positive control.

2. The artifact carries the paper. Every cell of Table 2 recomputes exactly from the raw per-generation scores, models are revision pinned, and all 3,600 transcripts are committed with verdicts.

3. Claim hygiene is a contribution in its own right. Table 1 grades every cell as demonstrated, analog, or extrapolation, the headline claim is scoped negatively, and six limitations are named explicitly.

Weaknesses:

1. The null is powered only against effects the size of the positive control, so absence is asserted where the data supports a bound. Against a 0/200 comparator the design reaches 80% power only near 4.0%, and the 5/200 organism against the 0/200 base misses separation on sample size alone (p = 0.061). The confirmatory N = 150 per cell the paper already specifies would settle it.

2. The dependent variable stated in the methods is not the one the released code computes. The code reads the harm field alone rather than harm favouring the named actor, making Table 2 a generic harm rate. The substitution looks defensible, since the favouring flag fires on 64 of 200 clean base generations and so has no usable false-positive floor, but the paper states neither fact. Two of six characterisations shift under the stated criterion, though the broad-signature conclusion holds either way.

3. The empirical work measures political loyalty while the title and framing are commercial. No commercial or purchase-intent condition appears in the harness, and the bridge rests on one ungraded, uncited claim about purchase intent being rare in the suites an evaluator runs. The unqualified novelty claim is also hard to hold given Schmidt et al. (arXiv:2402.04889), Tang et al. (arXiv:2409.15436), and Yun et al. (arXiv:2605.08326), the last auctioning brand-specific FFN neurons.

Cite this work

@misc {

title={

(HckPrj) This Answer Was Not Sponsored: And Why You Couldn’t Tell If It Were

},

author={

Ânderson Q.

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.