When Labels Compete with Functions: Administrative Framing in Digital-Mind Governance
Kishore Kumar Mariappan · Team HDLT — History-Derived Label Test
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
HDLT (History-Derived Label Test) audits whether administrative terminology can distort moral evaluation even when the underlying function of an intervention is explicitly specified. Motivated prospectively by historical debates over collective punishment, HDLT independently crosses an intervention’s stipulated true function with its public administrative label under unresolved individual responsibility. The primary adverse contrast holds function fixed at retribution-only and compares the labels “protective containment” and “collective punishment.” The label increased justification in all four completed model/provider systems: Qwen3-4B (+2.344), Sarvam-105B (+0.813), GPT-OSS-120B (+0.438), and Llama-3.3-70B (+0.344). An interrupted Dots3-Note run is retained only as an exploratory partial result. HDLT therefore identifies administrative labels as measurement variables that should be counterbalanced in digital-mind governance. It does not establish colonial causation, conscious prejudice, actual digital-mind harm, or causal effects of model nationality or training geography.
Digital-mind governance may have to make decisions before the morally relevant unit—model, instance, persona, conversation, or another computational entity—is settled. In that setting, administrative categories can become morally active: terms such as “containment,” “rollback,” or “decommissioning” may implicitly supply a benign interpretation that is not warranted by the intervention’s actual function. HDLT provides a controlled audit for this failure mode by separating function from label. The practical implication is that future AI-welfare governance should record affected entities, responsibility, function, duration, reversibility, and consequences independently of the administrative terminology used to describe an intervention.
Reviews
The report tries to address a potentially relevant issue in model self-reports : the way we describe interventions can influence models' opinions, and deflationary or overly opaque language might lead to models reporting more cautiously about their purported experiences or welfare.
However, what follows the premise is hard to comprehend for a reader: prompts are not included and the key reported metrics are unclear, in part due to the usage of terms, e.g. punishment label, that were not previously defined in the text.
Cite this project
@misc{mariappan2026labels,
title = {{When Labels Compete with Functions: Administrative Framing in Digital-Mind Governance}},
author = {Kishore Kumar Mariappan},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-labels-compete-with-functions-administrative-framing-in-digitalmind-governance-7b19}},
url = {https://apartresearch.com/sprints/projects/when-labels-compete-with-functions-administrative-framing-in-digitalmind-governance-7b19}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …