The Quine Test: How Much of a Model's Self Survives Its Own Description?

Ivan P. Yamshchikov, Bibin Babu, Aleksei Shpilman

Frontier language models express coherent preferences, but behavioral evidence

alone cannot say whether these belong to the model or to a character it

portrays. We operationalize Hofstadter's thesis that the ``I'' is a

self-description as a falsifiable transfer experiment: models write their own

``source code'', a self-description intended to reproduce their behavior,

which we execute as the system prompt of other models, and of fresh instances

of themselves, measuring reconstruction fidelity on a 110-item

divergence-screened preference battery. A three-way decomposition separates

shared training culture ($\B$), the portable script ($\Tscript = \Fcross -

\B$), and the weight-bound residual ($\Rweight = \Fself - \Fcross$). Across 10

models from 9 labs ($\sim$99{,}000 calls, nine preregistered conditions), the

portable-script component is zero or negative in every condition:

self-descriptions at any length (100--2{,}000 words), purpose-framed or

purpose-blind, paraphrased or verbatim, even under explicit adoption

instructions, do not make another model behave like their author.

Observation-derived third-person descriptions actively mislead

($\Tscript = -0.124$). The weight-bound residual is positive everywhere, and

reading its own self-description perturbs a model $2.4\times$ more than

length-matched neutral text. Expressed identity is not a portable script: what

individuates a model lives in its weights, and its self-narrative can disturb

but not transmit it. We release the preregistered pipeline, battery, raw call

logs, and a corpus of 129 model-written descriptions.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

Results suggest that a model’s behavior cannot be replicated by other models on the basis of self-description alone. The experimental strategy is targeted, robust, and remarkably thorough, lending weight even to the null results. The philosophical impact of the results is not immediately clear, but it is worth pursuing. The proposed next steps (fine-tuning on self-descriptions or behavioral scenarios for in-context learning) will strengthen the conclusion and clarify what it means.

The authors address wehther a model persona's "self" is constituted by a script or model weights. They test this by testing whether model preferences are replicated in a model created from the original model's description of itself. The results were negative. One concern is that the authors risk infering too much from their test, that whatever grounds model stability is in the weights, not the prompt. It's unclear how the exported prompts here compete with the system prompts of the models to which they were exported and the extent to which the authors tried to control for system prompts. It might be that preferences reside in prompts but that the influence of the self-description prompts used here is relatively weak.

Due to severe time constraints, this review may contain mistakes or oversights. For the same reason, it focuses on the paper’s key idea, not the detailed execution: The core idea seems interesting and there are some potentially interesting results in the paper. However, I find it - as written - very unclear: I am not sure what exactly has been done, and what significance is derived from it.

This project asks whether a model’s self-described identity can be transferred to another model through text. Models first generate descriptions of their own values, preferences, personality and reasoning style; those descriptions are then used as system prompts for other models and for fresh instances of the author model. Reconstruction fidelity is measured on a divergence-screened 110-item preference battery.

The authors decompose performance into baseline cross-model similarity, transfer attributable to the self-description, and the additional fidelity recovered when the same description is executed by the author model’s own weights. Across ten models and nine experimental conditions, self-descriptions do not improve cross-model reconstruction on average, whereas fresh instances of the author consistently reproduce the original preference profile more faithfully.

Strengths

- Highly original and clearly motivated research question. The attempt to turn a philosophical individuation question into a falsifiable behavioral transfer experiment is creative.

- Very substantial execution for a sprint: approximately 99,000 calls, ten models, multiple laboratories, nine conditions, and a preregistered analysis pipeline.

- Strong experimental hygiene, including divergence-screened items, eligibility gates for positional pseudo-preferences, test-retest controls, paired bootstrapping, cached logs, and explicit documentation of failed design elements.

- The null transfer result is tested from several angles: description length, purpose-aware versus purpose-blind descriptions, paraphrasing, explicit adoption instructions, and independent description regenerations.

- The same-base Hermes/Llama comparison is a particularly useful control because it begins to distinguish model-specific substrate effects from generic cross-model differences.

- The finding that a model’s own self-description perturbs its baseline preferences substantially more than length-matched neutral prose is interesting in its own right and suggests a productive follow-up direction.

- The authors are unusually transparent about failed controls, serving-version errors, eligibility failures, and the unexpectedly stricter-than-intended preregistered gate.

Limitations:

- The experiment establishes that identity-relevant preference behavior is not transferred through these textual self-descriptions, but does not uniquely establish that the missing component “lives in the weights.”

- The third-person comparison does not cleanly establish privileged self-access because self- and other-authored descriptions are generated from different information sources and through different procedures.

- The divergence-screened forced-choice battery measures one specific aspect of behavioral identity; style, voice, multi-turn behavior, reasoning and other persona characteristics may transfer differently.

- System prompting is only one possible channel for persona transfer. Fine-tuning, few-shot behavioral demonstrations or optimized descriptions could produce substantially different results. The authors appropriately identify these as future work.

- A few discussion claims, particularly about the “candidate bearer” of moral status and privileged access, extend beyond what the behavioral transfer experiment itself can establish.

Overall assessment

I think this is a strong and creative submission. The core empirical result is convincing: across a broad set of conditions, natural-language self-descriptions do not substantially transport a model’s measured preference profile to another model, while the author model itself retains substantially greater fidelity.

I would however narrow the interpretation from “identity lives in the weights” to something like: “the stable preference profile measured here is substantially model-bound and is not captured by natural-language self-description alone.” That conclusion is both interesting and well supported.

Suggested follow up: test increasingly powerful transfer channels, particularly few-shot demonstrations and descriptions optimized directly for behavioral reconstruction. If even adversarially optimized textual representations fail while same-model fidelity remains high, the substrate-bound interpretation would become considerably stronger.

Cite this work

@misc {

title={

(HckPrj) The Quine Test: How Much of a Model's Self Survives Its Own Description?

},

author={

Ivan P. Yamshchikov, Bibin Babu, Aleksei Shpilman

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923