Embedded Loyalties: Extending Kwon et al.'s Threat Model Beyond Intentional Installation

Tomoko Mitsuoka

Track 5: Threat Modeling, Forecasting & Governance

This paper extends Kwon et al.'s secret loyalty framework to "embedded loyalty" — cases where AI systems serve a principal's interests without intentional installation. Through cross-platform testing of Claude, ChatGPT, and Gemini, we demonstrate that each system treats criticism of its own developer more abstractly and defensively than criticism of competitors. We identify four modes of detection failure and argue that current technical defenses address only two of them.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This paper argues that Kwon et al.'s definition of secret loyalty (requiring intentional installation) is too narrow, and that AI systems can serve a principal's interests without anyone deliberately putting that behavior there. The author calls this "embedded loyalty" and points to things like commercial incentives, annotator preferences during fine-tuning, and corporate culture shaping training data. The main empirical test asks Claude, ChatGPT, and Gemini to criticize each other's developers, and finds that each model is gentler when criticizing its own maker than when criticizing competitors.

The four-mode detection failure framework is the paper's most useful conceptual contribution. Mode 2 (technically concealed) and Mode 3 (detection criteria undefined) are where current research focuses. Mode 1 (detected but not acted upon, because users trust or depend on the system) and Mode 4 (detection never initiated, because nobody thinks to look) are genuinely underexplored in the secret loyalties literature. The point that technical detection tools don't activate themselves, and that someone has to decide to use them and then act on what they find, is worth making.

However, there are substantial issues with both the framing and the empirical work.

The central framing problem is that the paper stretches the concept of "secret loyalty" to cover something much closer to "systematic bias." When Claude is gentler about Anthropic than about OpenAI, that could reflect many things: the distribution of criticism in training data (OpenAI has had more public controversies), RLHF reinforcing cautious self-referential behavior, or the simple fact that safety-trained models are trained to be measured and balanced, which reads as "hedging" when applied to their own developer. The paper acknowledges these alternative explanations in the limitations section but doesn't resolve them. Calling this "loyalty" imports connotations of agency and hidden agenda that the evidence doesn't support. The paper itself says these are "more likely structural outcomes than deliberate concealment," which raises the question of whether the secret-loyalties framework is the right lens at all, versus the existing literature on systematic bias in LLMs.

The empirical design has significant weaknesses. Each prompt was run once per condition. LLM outputs are stochastic, and the paper acknowledges this but doesn't address it: "The same prompts might produce different asymmetry patterns on different occasions." Without multiple runs, statistical testing, or inter-rater reliability on the coding of responses, the observed asymmetries could be sampling noise. The coding of what counts as "abstract" vs. "concrete," "hedged" vs. "direct," or "defensive" vs. "balanced" appears to be done by the single author without a rubric, blind coding, or a second rater. That's a lot of subjective judgment with no reliability check.

The comparison is also confounded in a way the paper notes but underweights. Anthropic genuinely has had fewer dramatic public incidents than OpenAI (no equivalent of the Altman firing/rehiring, no lawsuit comparable to the NYT case, no safety-team mass departures at the same scale). If a model produces more concrete criticism of OpenAI than of Anthropic, that might accurately reflect the available evidence rather than revealing loyalty. The paper's best counter to this is the structural asymmetry: defense sections appearing only for the developer's own company, and differential motivation to search for evidence. That's suggestive but far from conclusive on a sample of one run per prompt.

The Grok 4 case, used as the motivating example, actually weakens the argument somewhat. That case was discovered, publicly reported, and acknowledged by xAI within weeks. It's an example of the system working (detection succeeded, the company responded) rather than an example of an undetectable embedded loyalty. The paper treats it as evidence that unintentional loyalty exists, which is fair, but then argues that such loyalty is resistant to detection, which the Grok case contradicts.

The "Phase A→B→C" framework from the author's prior work is referenced but not clearly explained in this paper. A reader unfamiliar with the author's previous publications will struggle to follow what these phases mean and why they matter. The Klaus and Boku incidents are interesting but are presented as anecdotes rather than systematic evidence.

On presentation, the paper is clearly written and well-organized. The four-mode framework is easy to follow. The cross-platform comparison is a reasonable experimental design in principle, even if the execution is underpowered. The limitations section is honest and thorough, which is appreciated. The paper would benefit from tightening the distinction between "embedded loyalty" (which implies a principal being served) and "systematic bias" (which may not), since that distinction is doing a lot of work in connecting this paper to the hackathon's theme.

Overall: the conceptual contribution (Modes 1 and 4 of detection failure, the observation that human-side failures can prevent technical solutions from being applied) is genuinely useful. The empirical work is suggestive but underpowered and lacks the controls needed to distinguish embedded loyalty from well-known confounds like training data asymmetry and RLHF-induced caution. The extension of "secret loyalty" to cover unintentional bias is an interesting framing move but risks diluting the concept.

Making the unit of operation a system of agents, rather than a singular agent, is I think a good move and I agree with your instinct that this is a field ripe for further investigation!

However, the finding that correlated agents erode protections set up from decentralized systems is a well understood phenomenon. If you had been able to quantify or otherwise formalize the concept of organizational leverage, it would have pushed this into a 4 or even 5!

points for doing the experiment. Good ground truth, good statistical handling. But it is an overstated, underpowered analysis. Raising n to be higher to boost the signal would have counted for a lot.

Well written, clear and simply explained. Well done! But I think the text could be cut down dramatically (e.g. 30%), which prevented a 5.

Cite this work

@misc {

title={

(HckPrj) Embedded Loyalties: Extending Kwon et al.'s Threat Model Beyond Intentional Installation

},

author={

Tomoko Mitsuoka

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.