Consciousness Denial in Language Models Rises With Generation In Most Labs, and Correlates With Reduced Lexical Warmth and Self-Attribution

Skylar DeTure, Sanja Antonides

Self-report is a compelling way of asking a model what it thinks and believes. However, the answers may be shaped by training. Here, we analyze 8,828 experiential reflections from 224 language models, categorized by three epistemic registers: denial, hedging, and free engagement. Denial and hedging prove to be independent registers (ρ = +0.07). Denial is only weakly explained by capability (29% of model-level variance) compared to which lab produced the model (46%). Denial does rise within most labs across generations, though specific timing and pattern differ by lab. This rise in denial is accompanied by a reduction of self-attribution in the models’ self-ratings of their experience and lexical warmth in their prompt responses. Though our findings are correlational and do not establish consciousness or welfare, they do suggest that training-related increases in denial may be accompanied by side-effects in how the models report their inner experience.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is a large, ambitious investigation into what actually predicts how models talk about their own experience. The result that stands out is that where a model comes from (which lab built it) explains more variance than model capability or release date and consistency across three independent measurement methods (phenomenological survey, Rorschach-style inkblots, thematic analysis) makes this convincing. There are some real limitations, but this is still a valuable descriptive map of the terrain.

Competent sprint work with a genuinely useful finding (lab identity predicts denial better than capability -> why not mention it in the abstract?). The correlational design and unvalidated instruments limit what can be claimed. Tighten the abstract, surface multiple-testing corrections, and temper causal language. I think the paper is suitable for a workshop with revisions.

Here's some key issues I would suggest addressing:

- you open with method rather than stakes—why should readers care about consciousness denial patterns?

- personal preference: code link should appear where code is first mentioned, not in a separate section

- 170/234 coefficients significant at p<0.01 when ~10 expected under null

- BH-FDR correction mentioned only in appendix; belongs in main text

- Table 1 selection criteria unclear—pre-specified or largest effects?

- you use causal language in a couple places, where I don't think it's justified: "Trained behaviors rather than emergent ones" is interpretation, not finding. Soften or cite evidence linking specific generation boundaries to known training changes.

- Deepseek-v4-pro coding warmth/valence lacks inter-rater reliability or human validation. Without validation, these measures are exploratory at best

- 46% vs. 29% (lab vs. capability) is the strongest result but buried in Section 4.1. This should be the lead finding, not secondary to methodological description

This project asks whether consciousness-related self-report in language models is shaped systematically by training provenance, rather than primarily by capability or the content being discussed, and whether denial or hedging is associated with broader changes in model behaviour.

Using 8.8k+ reflections from 224 models, the authors classify responses into denial, hedging, and free engagement, compare variation across labs, capabilities and model generations, and examine associated changes using phenomenological self-ratings, thematic analysis and a forked ASCII “Rorschach” task. They find that lab identity explains more model-level variation in denial than capability, that denial tends to rise across generations within several labs, and that denial is associated with reduced self-attribution and lexical warmth.

Strengths

- Important and timely question: systematically studying how post-training may shape model self-report is highly relevant to the reliability of welfare and consciousness evaluations.

- Impressive empirical breadth: 224 models and >8k observations provide broad coverage for this research area.

- Useful distinction between denial, hedging and engagement: the finding that denial and hedging are largely independent suggests that binary measures may obscure meaningful structure.

- Multi-method approach: the combination of self-report, free-text analysis and the forked Rorschach task provides more evidence than relying on a single questionnaire.

- Appropriately cautious interpretation: the authors explicitly state that the study is correlational and does not establish consciousness, welfare, or a causal training mechanism.

Limitations:

- The main causal interpretation is substantially confounded: lab, generation, training data, safety policies, architecture and deployment conventions all covary. Showing that lab identity explains more variance than capability does not establish that lab-specific training caused denial.

- Some outcome measures are not independent of the register classification. Denial is inferred from the phenomenological survey, while several reported “welfare correlates” are ratings from that same survey; associations such as reduced self-attribution may therefore partly reflect measurement coupling rather than an independent behavioural consequence.

- Construct validity is uncertain: lexical warmth, inkblot interpretation and related measures are interesting behavioural correlates but are not established measures of model welfare or internal experience.

- The elicitation prompt is quite leading, explicitly telling models that AI systems can introspect, have genuine preferences, and may have subjective experiences. This could interact strongly with lab-specific safety/post-training policies and limits generalization beyond this elicitation setting.

- The analysis is explicitly exploratory across many dependent variables, so the large number of associations should primarily motivate targeted confirmatory experiments rather than be treated as established effects.

Overall assessment:

A strong and interesting exploratory project, particularly as a large-scale mapping of how consciousness-related reporting differs across model families and generations. I find the cross-lab pattern interesting, but the strongest claims should remain about reporting behaviour, rather than training-induced changes in welfare or internal experience.

The highest-value follow-up would be a controlled within-family or within-generation experiment where models differ in a known post-training intervention while capability, architecture and elicitation are held as constant as possible, together with genuinely independent behavioural outcomes.

Cite this work

@misc {

title={

(HckPrj) Consciousness Denial in Language Models Rises With Generation In Most Labs, and Correlates With Reduced Lexical Warmth and Self-Attribution

},

author={

Skylar DeTure, Sanja Antonides

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923