The Interface Is the Intervention: A Preregistered Multi-Model Audit of Persona-Framed Synthetic Triage
Frank Peterlein
Persona prompts are often treated as behavioral interventions, but the measurement interface can dominate what becomes observable. We preregistered a 192-call synthetic triage battery crossing four system prompts, two response formats, six pairwise items, four repetitions, and alternating A/B order, then extended it across four open models with a frozen smoke gate. Persona effects were small relative to wording, order-associated variation, response format, and schema compatibility. Allowing uncertainty eliminated directional responses in completed runs, and two models failed the unchanged schema gate. The result is a measurement audit: persona is measurable, but not dominant or invariant to the interface.
This is an unusually careful and useful measurement audit. The preregistered reference run, prospective multi-model extension, binding smoke gate, frozen parser, retained invalid outputs, wording floor, and explicit refusal to make population or clinical claims are exemplary. Treating schema failures and universal uncertainty responses as instrument outcomes rather than silently coercing them into A/B choices is exactly the right methodological instinct.
The main limitation is resolution. Persona contrasts are estimated from six fixed items and four repetitions, and only Qwen and Phi reached the persona phase. The increments are therefore coarse, and the order-associated contrast is confounded with stochastic replicate. It is reasonable to say order/replicate variation exceeded the observed persona excess descriptively, but the current design cannot determine how much was caused by option order versus sampling noise. The union-style schema example is also an avoidable instrument defect: it caused two of four models to stop before the substantive test, so those results primarily diagnose the schema rather than the models.
The proposed v0.2 is the right next step: use a valid non-priming schema, pair canonical and reversed order within matched seed blocks, independently author neutral and persona paraphrases, expand scenarios and repetitions, and decompose item, wording, order, and generation variance. Minimal persona manipulations would complement the current multi-feature stress tests. This project makes a strong contribution by showing that uncertainty and incompatibility are often properties of the observation channel, not evidence that a model lacks a preference.
Your epistemic discipline is the strongest part of this work. You preregistered the study before the first call, and you pinned the checkpoints. You defined a neutral-paraphrase wording floor, so that persona effects have a benchmark against ordinary rephrasing. You also let a binding smoke gate stop two models, instead of coercing their outputs. Two limits remain, one statistical and one at the design level. Six items with four repetitions leave the persona-excess estimates (0.083, 0.036) on very coarse rates. Your alternating-order scheme also confounds display order with the stochastic replicate. You can therefore say least about your largest observed contrast. A brief pre-run schema check also finds the union-string schema priming that stopped SmolLM3 and OLMo. Your own v0.2 plan is the right next step, with seed-paired order reversal, non-priming schemas, and more items. This framework deserves estimates that its design can support.
Cite this work
@misc {
title={
(HckPrj) The Interface Is the Intervention: A Preregistered Multi-Model Audit of Persona-Framed Synthetic Triage
},
author={
Frank Peterlein
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


