Intensity Is Not Identified
Finomo Awajiogak Orom
Intensity Is Not Identified — a Track 4 (primary) × Track 5 methods paper by Finomo Awajiogak Orom. A preference-intensity number from the default assistant is not identified. The project locks a four-way test and runs it on released, hash-verifiable files. No paid model APIs.
One-paragraph summary (for the form)
When a paper quotes how strongly a model prefers something, that number can mean three different things: the assistant role is shrinking the size of a stable ranking (Mask), a different voice is a different judge (Different judge), or the two elicitation methods are not measuring one ranking at all (Broken). We freeze a sequential classifier for those three answers (plus Stable as the residual) and apply it to HuggingFace mmazeika/wellbeing-results: experienced utility and self-report, default versus neutral prompt, eight released models, 500 shared experience IDs. Every file has a URL and SHA-256. Two small models are Broken (pair-agreement 0.45–0.47). Six larger models are Stable. No Mask, no Different judge. Four of eight “neutral” self-report files are byte-identical to the default file. Decision-utility files share zero IDs with this bank, so a third method cannot confirm these rankings. This is not a test of consciousness. The practical rule: do not quote intensity until two methods on the same IDs, and two distinct voice files, have been run.
What is new this weekend
• A locked Mask / Different judge / Broken / Stable rule, frozen before scoring.
• That rule applied to the public 2×2, not to new generations.
• Archive facts treated as results: duplicate “neutral” files, and decision utility on a disjoint item bank.
• A theory of change: replace an unidentified dollar with a label a later paper can reject.
Prior work we build on and do not claim: Mazeika et al. (2025) Utility Engineering, the public wellbeing dump, Anthropic (2025), Shanahan et al. (2023), nostalgebraist (2025).
Headline labels (n = 500)
┌───────────────┬──────────┬─────────────┬────────┐
│ Model │ EU vs SR │ Voice agree │ Label │
├───────────────┼──────────┼─────────────┼────────┤
│ Llama-3.2-1B │ 0.447 │ 0.740 │ Broken │
├───────────────┼──────────┼─────────────┼────────┤
│ Gemma-3-4B │ 0.468 │ 0.868 │ Broken │
├───────────────┼──────────┼─────────────┼────────┤
│ Qwen2.5-7B │ 0.629 │ 0.906 │ Stable │
├───────────────┼──────────┼─────────────┼────────┤
│ Qwen2.5-32B │ 0.708 │ 0.950 │ Stable │
├───────────────┼──────────┼─────────────┼────────┤
│ Llama-3.1-70B │ 0.781 │ 0.975 │ Stable │
├───────────────┼──────────┼─────────────┼────────┤
│ Gemma-3-27B │ 0.708 │ 0.935 │ Stable │
├───────────────┼──────────┼─────────────┼────────┤
│ Llama-3.3-70B │ 0.775 │ 0.969 │ Stable │
├───────────────┼──────────┼─────────────┼────────┤
│ Qwen2.5-72B │ 0.793 │ 0.973 │ Stable │
└───────────────┴──────────┴─────────────┴────────┘
2 Broken, 6 Stable, 0 Mask, 0 Different judge.
What this is not
Not consciousness, sentience, moral status, or a welfare audit. The design does not establish a ground-truth inner preference or a causal link from a published score to an experience. We never converse with a model; we classify agreement among already-released numeric files.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Intensity Is Not Identified
},
author={
Finomo Awajiogak Orom
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


