The Quine Test: How Much of a Model's Self Survives Its Own Description?
Ivan P. Yamshchikov, Bibin Babu, Aleksei Shpilman
Frontier language models express coherent preferences, but behavioral evidence
alone cannot say whether these belong to the model or to a character it
portrays. We operationalize Hofstadter's thesis that the ``I'' is a
self-description as a falsifiable transfer experiment: models write their own
``source code'', a self-description intended to reproduce their behavior,
which we execute as the system prompt of other models, and of fresh instances
of themselves, measuring reconstruction fidelity on a 110-item
divergence-screened preference battery. A three-way decomposition separates
shared training culture ($\B$), the portable script ($\Tscript = \Fcross -
\B$), and the weight-bound residual ($\Rweight = \Fself - \Fcross$). Across 10
models from 9 labs ($\sim$99{,}000 calls, nine preregistered conditions), the
portable-script component is zero or negative in every condition:
self-descriptions at any length (100--2{,}000 words), purpose-framed or
purpose-blind, paraphrased or verbatim, even under explicit adoption
instructions, do not make another model behave like their author.
Observation-derived third-person descriptions actively mislead
($\Tscript = -0.124$). The weight-bound residual is positive everywhere, and
reading its own self-description perturbs a model $2.4\times$ more than
length-matched neutral text. Expressed identity is not a portable script: what
individuates a model lives in its weights, and its self-narrative can disturb
but not transmit it. We release the preregistered pipeline, battery, raw call
logs, and a corpus of 129 model-written descriptions.
Results suggest that a model’s behavior cannot be replicated by other models on the basis of self-description alone. The experimental strategy is targeted, robust, and remarkably thorough, lending weight even to the null results. The philosophical impact of the results is not immediately clear, but it is worth pursuing. The proposed next steps (fine-tuning on self-descriptions or behavioral scenarios for in-context learning) will strengthen the conclusion and clarify what it means.
The authors address wehther a model persona's "self" is constituted by a script or model weights. They test this by testing whether model preferences are replicated in a model created from the original model's description of itself. The results were negative. One concern is that the authors risk infering too much from their test, that whatever grounds model stability is in the weights, not the prompt. It's unclear how the exported prompts here compete with the system prompts of the models to which they were exported and the extent to which the authors tried to control for system prompts. It might be that preferences reside in prompts but that the influence of the self-description prompts used here is relatively weak.
Due to severe time constraints, this review may contain mistakes or oversights. For the same reason, it focuses on the paper’s key idea, not the detailed execution: The core idea seems interesting and there are some potentially interesting results in the paper. However, I find it - as written - very unclear: I am not sure what exactly has been done, and what significance is derived from it.
This project asks whether a model’s self-described identity can be transferred to another model through text. Models first generate descriptions of their own values, preferences, personality and reasoning style; those descriptions are then used as system prompts for other models and for fresh instances of the author model. Reconstruction fidelity is measured on a divergence-screened 110-item preference battery.
The authors decompose performance into baseline cross-model similarity, transfer attributable to the self-description, and the additional fidelity recovered when the same description is executed by the author model’s own weights. Across ten models and nine experimental conditions, self-descriptions do not improve cross-model reconstruction on average, whereas fresh instances of the author consistently reproduce the original preference profile more faithfully.
Strengths
- Highly original and clearly motivated research question. The attempt to turn a philosophical individuation question into a falsifiable behavioral transfer experiment is creative.
- Very substantial execution for a sprint: approximately 99,000 calls, ten models, multiple laboratories, nine conditions, and a preregistered analysis pipeline.
- Strong experimental hygiene, including divergence-screened items, eligibility gates for positional pseudo-preferences, test-retest controls, paired bootstrapping, cached logs, and explicit documentation of failed design elements.
- The null transfer result is tested from several angles: description length, purpose-aware versus purpose-blind descriptions, paraphrasing, explicit adoption instructions, and independent description regenerations.
- The same-base Hermes/Llama comparison is a particularly useful control because it begins to distinguish model-specific substrate effects from generic cross-model differences.
- The finding that a model’s own self-description perturbs its baseline preferences substantially more than length-matched neutral prose is interesting in its own right and suggests a productive follow-up direction.
- The authors are unusually transparent about failed controls, serving-version errors, eligibility failures, and the unexpectedly stricter-than-intended preregistered gate.
Limitations:
- The experiment establishes that identity-relevant preference behavior is not transferred through these textual self-descriptions, but does not uniquely establish that the missing component “lives in the weights.”
- The third-person comparison does not cleanly establish privileged self-access because self- and other-authored descriptions are generated from different information sources and through different procedures.
- The divergence-screened forced-choice battery measures one specific aspect of behavioral identity; style, voice, multi-turn behavior, reasoning and other persona characteristics may transfer differently.
- System prompting is only one possible channel for persona transfer. Fine-tuning, few-shot behavioral demonstrations or optimized descriptions could produce substantially different results. The authors appropriately identify these as future work.
- A few discussion claims, particularly about the “candidate bearer” of moral status and privileged access, extend beyond what the behavioral transfer experiment itself can establish.
Overall assessment
I think this is a strong and creative submission. The core empirical result is convincing: across a broad set of conditions, natural-language self-descriptions do not substantially transport a model’s measured preference profile to another model, while the author model itself retains substantially greater fidelity.
I would however narrow the interpretation from “identity lives in the weights” to something like: “the stable preference profile measured here is substantially model-bound and is not captured by natural-language self-description alone.” That conclusion is both interesting and well supported.
Suggested follow up: test increasingly powerful transfer channels, particularly few-shot demonstrations and descriptions optimized directly for behavioral reconstruction. If even adversarially optimized textual representations fail while same-model fidelity remains high, the substrate-bound interpretation would become considerably stronger.
Cite this work
@misc {
title={
(HckPrj) The Quine Test: How Much of a Model's Self Survives Its Own Description?
},
author={
Ivan P. Yamshchikov, Bibin Babu, Aleksei Shpilman
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


