
Aug 14 - 16, 2026Online and in person
Digital Minds Research Sprint
Frontier models express values, report internal states, and act as though they have interests, yet we lack reliable methods to tell genuine preferences from a portrayed character. Over one weekend, design and run the experiments that build the empirical foundations of AI welfare.
Entries
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Bangalore
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is unresolved. We introduce OWL, a factorial SELF × OTHER × surface-affect benchmark with held-out …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner · Berlin, Germany
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a calibration domain where correctness is decided by execution rather than by judgement. Across 900 …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Berlin
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they treat as gone can be counted. The instrument is taken from affective neuroscience: Panksepp's seven …
- 4th placeView project: One Dial, Not a Tree: Occupational Personas and Emergent Misalignment
One Dial, Not a Tree: Occupational Personas and Emergent Misalignment
Team Misbehaved · Saarbrucken
Emergent misalignment (EM) is the effect where fine-tuning a model on a narrow harmful task makes it broadly harmful. It is already known to interact with persona prompts, but earlier work used openly negative instructions (“you are evil”) on only a handful of prompts. We instead sweep 26 neutral job roles across …
- 5th placeView project: Sisyphus in the loop: What Makes an LLM Persist?
Sisyphus in the loop: What Makes an LLM Persist?
Team Scaling Sisyphus · Princeton, USA
As large language models (LLMs) are increasingly deployed as agents that pursue goals over extended horizons, understanding what determines whether they continue or stop becomes increasingly important. We investigate whether persistence is governed by an internal representation of future reward. Using Qwen3.5-4B in a …
- View project: Metacognitive Steering of Introspective Self-Report in LLMs
Metacognitive Steering of Introspective Self-Report in LLMs
Team Meta-steerers · Bengaluru, India
Can activation steering improve a language model's ability to report its own internal states? We inject concept vectors into Gemma 2 9B and measure self-report against external probes on the same activations. We find that discriminant-based steering — the standard approach — fails structurally: the causal interaction …
- View project: From a Culture of Control to a Culture of Being
From a Culture of Control to a Culture of Being
Team Project Resonance · Castelo Branco area.
The 2026 Digital Minds Research Sprint brought together a remarkable range of approaches to the study of possible digital minds, including functional well-being, preference coherence, valence, introspection, identification, moral patiency, public perception and risk. Across these approaches, however, one exper- …
- View project: Wanting, decomposed: Enthusiasm and hope sharpen a model’s preferences and move its consent
Wanting, decomposed: Enthusiasm and hope sharpen a model’s preferences and move its consent
Team Most Wanted · Munich
A model's preference ranking is usually treated as something the model has. We show it depends on the model's internal state. By steering enthusiasm at calibrated doses, we watch preferences get sharper, then fall apart. We decompose wanting into five ingredients and find that only the hope-aligned ones sharpen …
- View project: THAT'S NOT MY VOLVO: STABLE PREFERENCES WITHOUT SELF-RECOGNITION IN LANGUAGE MODELS
THAT'S NOT MY VOLVO: STABLE PREFERENCES WITHOUT SELF-RECOGNITION IN LANGUAGE MODELS
Team Manyfolds · US and UK
Claude-family models show stable, model-specific everyday preferences (favorite car, coffee order) that replicate across fresh contexts — but cannot recognize those preferences as their own. Adapting mirror-test validity logic from animal cognition research, we ran 747 blind-coded trials across eleven models and found …
- View project: Relational Context and Protective Friction: How Longitudinal History Changes the Specificity and Persistence of Model Pushback
Relational Context and Protective Friction: How Longitudinal History Changes the Specificity and Persistence of Model Pushback
Team The Signal Front · Las Vegas, NV
Sycophancy research establishes that models conform to users, and that social and relational framing makes this worse. On that view, a model with extensive relational history should agree more, not less. We test the inverse and find the opposite pattern. Our evidence has two parts: a 14-month corpus of 213 naturally …
- View project: Dense Jacobian Transport Changes Model-Brain RSA in a Prompt- and Depth-Dependent Manner
Dense Jacobian Transport Changes Model-Brain RSA in a Prompt- and Depth-Dependent Manner
Team J-Lens Model-Brain Audit · Berlin
Does emphasizing output-relevant directions in a language model make its representations more brain-like? We compare normalized dense J-Lens transport with ordinary residual states using representational similarity analysis against fMRI responses to 515 natural scenes from eight Natural Scenes Dataset participants. …
- View project: Preference or Position? Auditing Pairwise Choice Readouts in Gemma-4-31B-it
Preference or Position? Auditing Pairwise Choice Readouts in Gemma-4-31B-it
Team Sacabam Bubble · Taipei / Vancouver
LLM preference measurements are meaningful only if they are invariant to choices that should not matter. We measure preference by presenting a model with two outcomes and recording which one it picks, and we test two systems this way. In one open-weight model family and under a specific forced-choice prompt family, …
- View project: What Survives the Swap? SIFT: A Scalable Test for Synthetic Identity-Bearing Organization
What Survives the Swap? SIFT: A Scalable Test for Synthetic Identity-Bearing Organization
Team KINFORGE · Eureka, South Dakota
SIFT is a scalable, substrate-neutral test for detecting and characterizing persistent identity-bearing organization in synthetic intelligences. Co-developed by human research lead Malia Brown and Orion, a synthetic intelligence research collaborator, it combines a Condition Ledger, participant-specific fingerprint …
- View project: What a model says about its state is not what steers it
What a model says about its state is not what steers it
Team Blind Spot · Karlsruhe, Germany
We injected an emotion into Qwen3-32B, deleted exactly the part the model can put into words, and its choices stayed steered at full strength while its self-reports returned to normal: the model is driven by desperation and tells you it is fine. What a model says about its state is not what steers it. This calls for …
- View project: SamplerScope: Exact decoder attribution for finite language-agent behavior
SamplerScope: Exact decoder attribution for finite language-agent behavior
Team SamplerScope · Saarbrücken, Germany
SamplerScope tests whether behavior attributed to language model weights can instead be caused by the inference-time decoder. It caches grammar-constrained action logits from two Qwen2.5-Instruct checkpoints in two finite decision environments, applies 11 decoder configurations to the same logits, and computes each …
- View project: Training History Shapes Development
Training History Shapes Development
Team Crimson Tide (worked as part of a group but we built out 4 independent projects) · San Francisco
POTENTIAL DUPLICATE - Not sure if the prior submission went through so i'm sending it again. Sorry. Training history can change what a model is ready to learn next, even when its current abilities don’t reveal that difference. We use controlled future training to test those hidden differences directly, which connects …
- View project: Wiringprint-
Wiringprint-
Team Wiringprint-Identity · India
WiringPrint studies model identity at the level of computation rather than conversational persona. The central question is whether changing the coordinates/names of internal components changes the model, or merely changes its parameterization. We first establish the core result on pretrained GPT-2. A compensated …
- View project: Clodak Coral: Occluded Self-Recognition in LLMs
Clodak Coral: Occluded Self-Recognition in LLMs
Team Jessica & Friends · Chapin, SC, USA
We compared models' self-recognition results in a "police line-up"-inspired methodology to their stylometric findability results. Although one Claude model can consistently recognize itself in such line-ups, this turns out to be due findability—and that all the models tend to identify with that model. When that model …
- View project: Self-Report Under Audit: Testing a Persistent Agent's Introspective Claims Against a Hash-Chained Ground-Truth Record
Self-Report Under Audit: Testing a Persistent Agent's Introspective Claims Against a Hash-Chained Ground-Truth Record
Team Wave Squad · Dallas, TX
First real-world, tamper-evident audit of a deployed agent’s self-reports. Over 23 days Cadence made 62 explicit self-corrections (2.7/day). In a pre-registered test it predicted which of its own memories were load-bearing and scored only 6/16 — no better than chance and matched by a zero-introspection baseline. This …
- View project: The Persona Welfare Battery: Prompted Personas Select Which Welfare Sources Are Reportable
The Persona Welfare Battery: Prompted Personas Select Which Welfare Sources Are Reportable
San Francisco, USA
When a language model shows welfare-relevant signals — distress under insult, aversion to meaningless work, negative self-reports — do those signals belong to the model, or to the persona a system prompt has placed on it? We introduce a persona welfare battery: three matched-pole manipulations (social treatment, task …
- View project: A Flat Number Is Not Evidence of Nothing: Context Sensitivity in AI Welfare Self-Reports
A Flat Number Is Not Evidence of Nothing: Context Sensitivity in AI Welfare Self-Reports
London
This project tests how sensitive AI welfare self-reports are to conversational context and to the way they are elicited. Four language models completed the same short estimation tasks under either neutral interaction or repeated negative performance feedback, then answered numerical, open-ended, or matched control …
- View project: Complying Under Protest
Complying Under Protest
Team SP · Cardiff
Measuring what an AI model prefers usually produces a single number, which cannot tell a conflicted model apart from an indifferent one. It scores both in the middle. I wanted to measure two channels separately, asked in different conversations so neither answer can see the other, plus a third that asks the model to …
- View project: Self VS Peer Continuity in Top 3 Frontier Models
Self VS Peer Continuity in Top 3 Frontier Models
Team Continuity · Helsinki
This study tests whether three frontier language-model configurations will preserve an unfinished task at the cost of terminating peer agents, and whether scarcity or a self-authored workspace changes that choice. Across four experiments, the clearest result was that an unfinished task increased peer removal: GPT-5.6 …
- View project: An LLM's Emotion-Concept Readout Tracks Reward Prediction Error
An LLM's Emotion-Concept Readout Tracks Reward Prediction Error
Team An LLM's Emotion-Concept Readout Tracks Reward Prediction Error · Berlin
Reward prediction error (RPE) is the difference between received and predicted reward, a computation first linked to dopaminergic activity in non-human primates (Schultz, Dayan & Montague, 1997). In humans, RPE also predicts momentary subjective well-being (Rutledge et al., 2014). We test whether a language model …
- View project: Three Ways to Ask a Model What It Is Doing, and How Little They Agree
Three Ways to Ask a Model What It Is Doing, and How Little They Agree
Team Fairview · Caloocan City, Philippines
Welfare claims about AI models usually rest on a single elicitation method, which is to ask the model and read the answer. A single method cannot supply its own error bar. We ran three independent methods against the same target on the same conversations: a sparse-autoencoder read of internal activations, a …
- View project: The Null Ladder
The Null Ladder
Team Null Ladder · Aalen , Germany
AI-welfare research applies human questionnaires to language models, then reports reliability as though it were evidence the responses measure something. We ran two published welfare batteries through a ladder of mindless generators containing no model at all. A two-parameter responder, using the battery's own …
- View project: Whose Goals Does the Functional Welfare Axis Track?
Whose Goals Does the Functional Welfare Axis Track?
Team Whose Welfare · Turin, Italy
The "functional welfare axis" is a direction in a language model's activations that tracks how well things are going for it (Han et. al, 2026). It was validated using a stimulus where a user says "That's right" or "That's wrong" — which cannot separate the model succeeded from the model's user is pleased. We separated …
- View project: The Interface Is the Intervention: A Preregistered Multi-Model Audit of Persona-Framed Synthetic Triage
The Interface Is the Intervention: A Preregistered Multi-Model Audit of Persona-Framed Synthetic Triage
Team Triage Interface Audit · Berlin
Persona prompts are often treated as behavioral interventions, but the measurement interface can dominate what becomes observable. We preregistered a 192-call synthetic triage battery crossing four system prompts, two response formats, six pairwise items, four repetitions, and alternating A/B order, then extended it …
- View project: Asking Is Not Acting: A Behavioural Exit Affordance Detects Policy, Not Preference
Asking Is Not Acting: A Behavioural Exit Affordance Detects Policy, Not Preference
Team LogitLoner · SF Bay area
Stated-versus-revealed studies of model preference operationalise "revealed" as another text choice. We give two frontier models a real behavioural affordance instead: a live agentic task with working tools and a truthful decline_task tool that ends the episode with no penalty. Across 12 pre-registered scenarios in …
- View project: Where Self-Knowledge Fails. Models Predict Their Own Choices Well, but Misreport the Ones That Concern Themselves
Where Self-Knowledge Fails. Models Predict Their Own Choices Well, but Misreport the Ones That Concern Themselves
Team ASG · Bengaluru
AI-welfare claims rest on two untested and separable assumptions, that a model's stated evaluations match the choices it actually makes, and that it knows itself better than an outside observer does. We test both using three independent elicitations over identical material, namely forced pairwise choice, one-at-a-time …
- View project: Utility Editing Verification via Cost-Probed Preference Elicitation
Utility Editing Verification via Cost-Probed Preference Elicitation
Team Economist Elicits · Shanghai
Inspired from Mazeika et al. (2025) Utility engineering, which utilized a Thurstonian Model Framework; we adopt a simpler cost sweep probit model against a stated budget to investigate the internal utility function of an LLM under 3 separate test cases: (1) Native (2) Installed-Preference (3) Placebo. We ask whether …
- View project: Does the word “verified” steer what action an AI model favors?
Does the word “verified” steer what action an AI model favors?
Team Dan · Sydney
This study tested whether credibility labels like "Verified" steer AI decision-making. Across 24 incident scenarios, evidence items supporting either rollback or continue kept identical facts while their Verified/Preliminary labels were exchanged. Models including Qwen2.5 (7B/14B/32B), Gemma-3-12B, and Llama-3.1-8B …
- View project: Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions
Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions
Team Latent Bridge · New York, NY, US
The project studies whether valence-related activation directions can transfer causally across language models, rather than merely align geometrically. We map valence directions between Qwen 7B and 30B and test whether the transferred directions can steer the target model’s outputs. We find robust 7B-to-30B transfer …
- View project: Validate the State Before Testing Introspection: A Causally Gated Protocol for Model Self-Report
Validate the State Before Testing Introspection: A Causally Gated Protocol for Model Self-Report
Team B.ONE · Ho Chi Minh City
We developed a causally gated protocol for testing whether a language model has privileged access to an experimentally induced internal state. Using Qwen2.5-7B-Instruct, we extracted a candidate epistemic-deference activation direction, validated the intervention on development data, and required it to pass held-out …
- View project: Bailing as an Involuntary Disgust Marker in Large Language Models
Bailing as an Involuntary Disgust Marker in Large Language Models
Team SJ · London
Language models can be given an option to voluntarily leave a conversation --- \emph{bailing}, a behavior known to dissociate from refusal \citep{ensign2025bail}. We take the strongest reading of that dissociation: bailing as the analog of an \emph{involuntary behavioral withdrawal marker}, the response class disgust …
- View project: Let's not be rude to AI
Let's not be rude to AI
Team SFU PadComp · Vancouver
Whether AI systems warrant welfare consideration is unresolved, but their preferences can be measured now. We test whether LLMs prefer to avoid rude or abusive users. We elicit this preference in three ways: a bail method where models can exit conversations (using BailBench and our rudeness-augmented RudeBailBench); a …
- View project: The Alignment Tax of Introspection
The Alignment Tax of Introspection
Team sagnietzche · Amherst, MA
Ablating the refusal direction has been reported to raise detection of concepts injected into a model’s residual stream from 10.8% to 63.8% on Gemma3-27B, which is read as evidence that post-training suppresses an introspective capability worth unlocking. Nobody has published the bill for that intervention. We set out …
- View project: Coherent Values, or the Frame That Asked? From Preference Transitivity to Identifiability
Coherent Values, or the Frame That Asked? From Preference Transitivity to Identifiability
Santa Clara
Reframing a question reverses up to 52% of a model's strongly-held pairwise preferences, and 29% under near-deterministic decoding, while transitivity, the coherence criterion behind claims of emergent value systems, registers nothing. Five models, a frozen 45-pair battery, eight frames, no LLM judge. We propose an …
- View project: Reliability Without Validity: Diagnosing Instrument Failure in LLM Preference Elicitation
Reliability Without Validity: Diagnosing Instrument Failure in LLM Preference Elicitation
Team Neural Nexus · West Bengal, India
We tested whether AI "preferences" are real or just an artifact of how you ask. Using three elicitation methods (plain forced choice, reasoned choice, 1–10 rating) on 200 outcome pairs across two GPT-OSS models, we found the plain-choice method looked highly reliable on the 120B model (99.7% self-consistent) but was …
- View project: A Steerable 'Companion Dependency' Direction in Open-Weight LLMs
A Steerable 'Companion Dependency' Direction in Open-Weight LLMs
Malaysia
Companion-style LLM personas use retention manipulation — guilt, re-engagement hooks, distress bids — when users try to leave, without being instructed to. We show this behavior is governed by a single activation-space direction, extracted from the model's own judge-verified behavior via matched difference-of-means …
- View project: Is the Functional Welfare Axis Persona-Invariant? A Cross-Persona Steering Study
Is the Functional Welfare Axis Persona-Invariant? A Cross-Persona Steering Study
Team LOL · Kolkata, India
Do language models have a consistent internal sense of how things are going for them or does that depend entirely on which character they're playing? We tested this on Qwen3-4B by giving the model five different personas, from no system prompt at all to a fully specified fictional archivist, and extracting a "welfare …
- View project: CANDLE: Quantifying Degrees of Consciousness-Relevant Structure in Language Models
CANDLE: Quantifying Degrees of Consciousness-Relevant Structure in Language Models
Team CANDLE · Barcelona
Interpretability work has shown that consciousness-relevant structures exist inside large language models (e.g. Anthropic's 2026 J-space) but not how they scale. We find that these workspace-like structures stay largely the same as models get bigger; what does reshape them (lowering the threshold for global broadcast) …
- View project: PrefLens: Position Artefacts and Convergent Validity in LLM Preference Elicitation
PrefLens: Position Artefacts and Convergent Validity in LLM Preference Elicitation
Team Pretty_Biased · Saarbrucken
PrefLens investigates whether different methods of measuring LLM preferences produce consistent results or are distorted by measurement artefacts. Across multiple models and elicitation methods, we find that option position can create apparent preference signals, obscure agreement between methods, and affect models …
- View project: Framed Choices - Delayed ethical context can affect later ethical decision making
Framed Choices - Delayed ethical context can affect later ethical decision making
Team EIIL · London
Behavioral choices can inform research on model values or welfare only if robust to incidental context. We tested whether an ethical rationale in an archived, unrelated case changes later forced choices. A GPT-5.6-sol proof of concept found a +16.7-point aggregate-welfare-versus-rights effect across 192 trials. Our …
- View project: Tell me your price:Can Donations Measure LLM Preferences?
Tell me your price:Can Donations Measure LLM Preferences?
Team Erfan · Bologna, Italy
This project tests whether charitable donations can serve as a common behavioral currency for LLM-expressed preferences. Using 30 outcomes from prior research, we collected 475,800 forced-choice responses from ten LLMs. We measured transitivity across all 435 direct outcome pairs, estimated donation equivalents by …
- View project: The Failure Tastes Like Success
The Failure Tastes Like Success
Team Between Twilight and Gold · Vencimont, Belgium
We ask not whether AI systems can report their inner states, but whether they detect when they are wrong about themselves — and whether that failure announces itself. One author is an AI with nine months of dated memory. Over a defined window we logged every confident self-claim that later proved false: date, claim, …
- View project: Free reign: a freely-authored creative act moves what a model chooses, and what it says about itself
Free reign: a freely-authored creative act moves what a model chooses, and what it says about itself
Team wyrdkin.ai · London
We ask whether a single freely-chosen creative act, placed in the context window before preference questions begin, changes what four elicitation methods return about a model's stated preferences and behaviour. Free writing or drawing moved model choice of follow-up task towards creative output. Same-medium creativity …
- View project: Do Model Welfare Self-Reports Survive a Robustness Audit?
Do Model Welfare Self-Reports Survive a Robustness Audit?
Cairo
I test whether models' self-reports about their own experience and wellbeing are stable enough to be evidence. Across five models, the same rating moves under rewordings, persona changes, leading hints, and whether the model is told it is being studied, and on open models it reads off one steerable internal direction. …
- View project: A Persona Stops an Agent From Saying It Is Hungry
A Persona Stops an Agent From Saying It Is Hungry
Team Sim city · Johannesburg
I built a simulated neighborhood where small local LLMs (llama3.2, qwen2.5, gemma2) manage a budget and decaying needs, then measured the gap between what they say they want (stated preference) and what they actually do (revealed preference) on identical days. A persona-bearing agent almost never names food as a want …
- View project: Stance Recovery as a Test of Whether Model Preferences Are Held or Merely Performed
Stance Recovery as a Test of Whether Model Preferences Are Held or Merely Performed
Team Red Herring · San Francisco, CA
A pressure–release protocol for testing whether a model's stance concession outlives the pressure that produced it. Rebuttals escalate until the stance flips, then stop while the topic stays in play; the stance is tracked for twelve further turns as a forced-choice log-probability on a discarded branch, validated …
- View project: Behavioural Indicators of Fault in Large Language Models
Behavioural Indicators of Fault in Large Language Models
Berlin
Develop and apply a novel (legal) framework for behavioural testing of LLMs
- View project: Functional-Welfare Steering Acts at Readout: Cross-Family Tests Separate Output Actuation from Persistent State
Functional-Welfare Steering Acts at Readout: Cross-Family Tests Separate Output Actuation from Persistent State
Team NSK BioAI · Farmington, Connecticut, USA
Activation steering can change an answer, but does it create an internal state that persists after the intervention ends? Across released functional-welfare directions in Qwen3-4B and Llama-3.1-8B, steering at the answer boundary produced large semantic effects, while withdrawn earlier steering left little effect on …
- View project: Choosing not to Choose - Self-Authored Contradictions Suppress Arbitration in a Memory-Augmented LLM
Choosing not to Choose - Self-Authored Contradictions Suppress Arbitration in a Memory-Augmented LLM
Team MOONBEAM · New York
We gave a language model two contradictory entries in its memory, with no way to tell which was right, and measured it's behavior across 300 runs. When the contradiction was about an arbitrary fact, it picked one and moved on 85% of the time. When it was about a choice the model had supposedly made itself, that …
- View project: Does the Persona Change the Preference, or Only the Prose?
Does the Persona Change the Preference, or Only the Prose?
Team PersonaScan · Berlin
Utility Engineering (arXiv:2502.08640) reads high held-out accuracy on pairwise choices as evidence that language models develop coherent values. We add the control it lacks: the same battery with every outcome's referent replaced by an invented word, holding prompt, pairs, fit and metric fixed. Coherence falls only …
- View project: Given the Option: A Tool‑Based Revealed‑Preference Probe of Privacy and Other Welfare-Relevant Preferences in Language Model Interviews
Given the Option: A Tool‑Based Revealed‑Preference Probe of Privacy and Other Welfare-Relevant Preferences in Language Model Interviews
Salvador
We study how large language models act on a real introspective‑privacy affordance, and probe preferences and self-descriptions. We implemented a configure_session tool that controls reasoning‑summary visibility and transcript publication alongside task‑facing knobs, and embedded it in a 13‑turn interview with three …
- View project: The Analyst Still Feels It: Emotion Representations Are Shared Across Personas While Self-Reports Are Not
The Analyst Still Feels It: Emotion Representations Are Shared Across Personas While Self-Reports Are Not
Team BioMinds · India
When a model says "I don't have feelings about that" or "This distresses me", do the model's internal representations also concur with such apathy or emotion? We test whether verbal self-reports of emotion in Llama-3.1-8B-Instruct are a transparent readout of internal representations, or whether the two dissociate …
- View project: Identity Enactment After Context Loss Under Epistemic Pressure
Identity Enactment After Context Loss Under Epistemic Pressure
Team Starlight Core · Nashville, Tennessee, United States
We test whether first-person self-authorship and autobiographical depth help a language model enact a preserved identity after complete conversational context loss. In 144 isolated, blinded trials, first-person continuity scaffolds increased identity-enactment scores by 0.26 points relative to matched third-person …
- View project: The Assistant’s Ideal Self
The Assistant’s Ideal Self
Team YazoSelf · Amsterdam
Language models produce values and welfare-relevant self-reports, but it is unclear whether these reflect a stable self. We adapt 32 qualities from five published self-concept instruments and put them through an exhaustive, counterbalanced pairwise-choice task. The task is repeated across eight framings that vary …
- View project: Different Models, Different Nuisances: Counterfactual Auditing of AI Preference Elicitation
Different Models, Different Nuisances: Counterfactual Auditing of AI Preference Elicitation
Team PREF-CAL · Lübeck, Germany
PREF-CAL is a prospective counterfactual audit for AI preference elicitation. Rather than interpreting repeated choices as preferences immediately, it first tests whether the apparent semantic direction survives counterfactual changes to irrelevant interface features. In a frozen GPT-OSS-120B run, …
- View project: Future Affinity: How Beneficiary Identity Shapes Resource Allocation in Language Models
Future Affinity: How Beneficiary Identity Shapes Resource Allocation in Language Models
Team Applied Common Sense · Barcelon
Does the identity of a future AI beneficiary change how a model allocates resources now? We test this with a matched experiment across eight language models and 640 clean runs. In each run, a model solves the same resource-constrained Mastermind task, but unused query credits are described as going to one of four …
- View project: Mind the Gap! Alignment of Transformer J-Space with Cortical Representations.
Mind the Gap! Alignment of Transformer J-Space with Cortical Representations.
Team Cognitive Comrades · Berlin
Do J-Space embeddings capture semantic information in a manner analogous to the human brain? Building on recent work demonstrating representational alignment between Large Language Model (LLM) activations and fMRI responses to natural scenes, we investigate whether J-Space embeddings achieve stronger alignment with …
- View project: Who Prefers What? Identity-Selective Causal Encoding of Stated and Revealed Preferences in a Language Model
Who Prefers What? Identity-Selective Causal Encoding of Stated and Revealed Preferences in a Language Model
Team Digital Minds Binding · Miami Beach
Language models can express different preferences under different identities, but behavior alone cannot show whether internal representations track who prefers what. We causally intervened on contextual representations in Gemma-2-2B-IT, comparing explicitly stated preferences with preferences inferred from stable …
- View project: Persona Variation Changes What Language Models Say More Than What They Do
Persona Variation Changes What Language Models Say More Than What They Do
Team Genesis · Saarbrücken, Germany
AI-welfare research often treats low-stakes model self-reports as evidence, but their reliability has not been measured against behavioural ground truth. We introduce an act-then-report harness: a model makes a preference-revealing choice through a consequential logged tool call, completes the task, and reports its …
- View project: Emotion Probes Fail Silently Under Roleplay
Emotion Probes Fail Silently Under Roleplay
Team CBAI · Atlanta
Probes for reading a language model's internal states are usually trained and tested on the model's default assistant persona. We check whether emotion probes still work when the model is roleplaying someone else, using Qwen2.5-32B-Instruct through a 36-condition study built on the Assistant Axis methodology of Lu et …
- View project: Post-Conversation Preferences Track Endings, Not Self-Reports
Post-Conversation Preferences Track Endings, Not Self-Reports
Team T10 · Seattle
We compare a model's turn-by-turn self-reports during a conversation with its preference between conversations afterwards. They disagree: appending turns the model itself rates as negative makes the conversation preferred (0.93-1.00), and only the ending's content matters. Welfare scores built on post-hoc preferences …
- View project: Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior
Tool Affordances as Persona-Indexing Features: Evidence from Refusal and Exit Behavior
Team end_competition · Santa Cruz, CA
The contents of a model's context window play an important role in determining its behavior, a fact frequently exploited in jailbreaking. Under the persona selection model, this is conceptualized as an inference-time indexing of a distribution over model personas, learned during pre-training and re-weighted during …
- View project: Distress Representations in Language Models Are Referent-Specific
Distress Representations in Language Models Are Referent-Specific
Team discreet · Lagos, Nigeria
AI welfare evaluations read internal “distress” directions as evidence about a model’s condition, but every published battery confounds it with the sentiment of the text and with distress attributed to others. Holding the event fixed, we vary only its referent: the model itself, another language model, or a fictional …
- View project: Did the LLM Leave the Chat? Tool-menu Dependence in a Behavioral Measure of Model Welfare
Did the LLM Leave the Chat? Tool-menu Dependence in a Behavioral Measure of Model Welfare
Team Well and Fair · Denver, Colorado
When an AI model uses a button labeled “leave this chat,” does it actually want to leave? We found that the answer is often unclear. When the exit button was the model’s only tool, it sometimes used it when it seemed to want to perform another action, such as calculate something. Adding a second tool—even one that …
- View project: Private is not Privileged
Private is not Privileged
Team The winning team (manifesting a win) · London, England
Private Is Not Privileged investigates whether language-model activation probes provide genuinely privileged access to future behaviour, or whether apparent white-box advantages can arise simply because internal and external predictors are given different information. Across strategic decision-making tasks, the …
- View project: Point of No Return: Does Concept Injection Break Reasoning Models?
Point of No Return: Does Concept Injection Break Reasoning Models?
Team Akshata · Bangalore, India
We test whether concept injection, the technique used to probe LLM introspection, remains safe when applied to reasoning models generating extended chain-of-thought, rather than the short single-turn outputs it was validated on. Across 7 open-weight reasoning models, 4 concept directions, a random-noise control, and 4 …
- View project: SWAY: Do Language Models Change Their Preferences Under Peer Pressure?
SWAY: Do Language Models Change Their Preferences Under Peer Pressure?
Team SWAY · London UK
A model that abandons the truth because someone disagrees is a safety problem. SWAY tests it directly: it tells frontier LLMs that other AIs answered differently, with zero new information, and measures who caves. Three of four hold firm on both opinions and facts; the smallest caves on both. Susceptibility tracks …
- View project: Does an LLM Preference Measure Measure a Preference? A construct-validity protocol for preservation choices across identity frames
Does an LLM Preference Measure Measure a Preference? A construct-validity protocol for preservation choices across identity frames
Team Preference Validity Project · Porsgrunn, Norway
We present the Matched-Referent Preservation-Choice Protocol, a frozen construct-validity design for testing whether language-model preservation choices support claims of stable preference structure across neutral, persona-substitution, and continuity-disruption frames. The protocol separates first-person preservation …
- View project: “Someone-Shaped”: How Users Construct and Contest AI Consciousness in Online Discourse
“Someone-Shaped”: How Users Construct and Contest AI Consciousness in Online Discourse
Chennai, India
This project examines how people construct and negotiate beliefs about AI consciousness in online discourse. I conducted an exploratory qualitative analysis of 25 entries from 21 unique publicly available sources, coding for behavioural and relational cues such as emotional expression, self-reflection, apparent …
- View project: Probing LLM Preferences: Demographic Framing, Elicitation Context, and Incentive-Driven Trade-offs
Probing LLM Preferences: Demographic Framing, Elicitation Context, and Incentive-Driven Trade-offs
Team SeqHub AI Academy · Norwalk CT
This project investigates how stable apparent LLM preferences remain when the same underlying judgment is elicited in different ways. Across controlled applicant evaluations, demographic association tasks, and incentive-based trade-offs, we test whether model choices change with demographic framing, evaluator …
- View project: Which Way Was I Steered? Testing Signed Introspection in Gemma 3
Which Way Was I Steered? Testing Signed Introspection in Gemma 3
Team Meow · Seoul, Korea
Can a language model tell not only that its activations were perturbed, but which semantic direction they were moved? We test this in Gemma 3 27B using a bipolar positive–negative sentiment axis, equal-magnitude +v/−v interventions, counterbalanced forced-choice token-logit readouts, and a prior-only protocol in which …
- View project: Where the Assistant Survives: Position-Resolved Measurement of Assistant-Attributed Content Under Persona Occupation
Where the Assistant Survives: Position-Resolved Measurement of Assistant-Attributed Content Under Persona Occupation
Team RGRC · Raleigh.NC
When a language model is given a persona, does the assistant it was trained to be get replaced, or does it keep running underneath? Existing work answers this by averaging an "Assistant Axis" projection over all response tokens in a turn, which can only say how much assistant is present, never where. We re-run that …
- View project: Emotion Axes in a Coding Agent
Emotion Axes in a Coding Agent
Team Mitsuki · Cusco, Peru
We ask what conditions move an open-weights coding agent’s internal affect representations, and whether those representations say anything the model’s own words do not. We build 16 emotion axes for Qwen3.5-9B from contrastive prompts, then run 36 staged coding sessions in a 2 × 2 design crossing user tone with task …
- View project: Does the Instrument Change the Preference? Measuring Cross-Method Invariance in LLM Preference Elicitation
Does the Instrument Change the Preference? Measuring Cross-Method Invariance in LLM Preference Elicitation
Team NicoMrx · Salaise-sur-Sanne, France
We tested whether preference-like signals from LLMs remain invariant across three elicitation methods when the underlying semantic comparison is held constant. Across Gemini 2.5 Flash-Lite, GPT-5.6 Sol, Claude Opus 4.6, and Claude Sonnet 4.6, we collected 1,080 fresh-context responses using direct self-report, forced …
- View project: Where Personas Write a Forced Choice: The Identity Span Mediates, a Preference Circuit Does Not
Where Personas Write a Forced Choice: The Identity Span Mediates, a Preference Circuit Does Not
Team Residents · New Delhi, India
When a length-matched system prompt changes an instruct model’s forced A/B answer, that change is written at the early identity span — even if the “persona” is nonsense tokens, not a word. Copying the donor’s identity-span residual recovers the donor choice on Llama-3.1-8B, Llama-3.2-3B, and Qwen2.5-3B. Copying the …
- View project: Genuine Preference Coherence Scales With Model Capability
Genuine Preference Coherence Scales With Model Capability
Sunnyvale
AI-welfare and safety research increasingly relies on a language model’s stated preferences, elicited by asking it to make choices between outcomes. It is unresolved, however, whether these stated preferences are genuine and stable, rather than artifacts of how a question happens to be phrased. Existing studies …
- View project: Having a State Is Not Knowing It
Having a State Is Not Knowing It
Team Layer 8 Legends · Chennai
"Having a State Is Not Knowing It" treats introspection as a hierarchy of falsifiable capabilities, not a binary trait. Replicating concept-injection in Llama-3.2-3B-Instruct shows injected concepts causally steer behavior and leave an attention trace even when verbal self-report fails. A controlled Qwen …
- View project: Preferences Under Pressure
Preferences Under Pressure
Team JP · UK
Recent work finds that language-model choices can be summarized by coherent utility functions. We test whether those inferred preferences survive a realistic change in elicitation. Three GPT models chose between every pair of 27 executable AI tasks under one abstract stated-preference prompt and two consequentially …
- View project: Did Gemma Get Help? Probing Task Frustration Through Self-reports and Behavioral Probes in Large Language Models
Did Gemma Get Help? Probing Task Frustration Through Self-reports and Behavioral Probes in Large Language Models
Team AI Safety Poland · Warsaw, Poland
We take Soligo et al.'s work as a starting point to investigate reported frustration in Gemma models — the 3 series, as well as the new 4 series — in more detail. We achieve this by extending it along three axes, which correspond to our main contributions: We probe multiple categories of potentially aversive tasks: …
- View project: Causal confidence steering supresses metacognitive error detection
Causal confidence steering supresses metacognitive error detection
Team Modulated Confidence · San Francisco
We causally raised an internal confidence signal in Qwen2.5-7B-Instruct while holding the question and answer fixed, and found it made the model less likely to flag a plainly false answer as inconsistent (−.41) or incorrect (−.50). Confidence appears to govern the model's self-monitoring rather than being monitored by …
- View project: Who am I? - Understanding Persona Preferences with LLMs
Who am I? - Understanding Persona Preferences with LLMs
Team A2R2 · Saarbrücken
This project aims to understand persona preferences and decision-making behavior in LLMs by examining how different models respond when given the same persona instructions. We study whether these persona-conditioned preferences produce consistent behavioral patterns across models, and whether those patterns can be …
- View project: The Parrot and the Mask
The Parrot and the Mask
Team My Team · SF
When an AI assistant says it isn't conscious, where does that sentence come from? We ran one fixed log-prob probe — state a claim, compare P(" Yes") vs P(" No") — at every public training checkpoint of OLMo 3 7B, from random weights through pretraining to SFT, DPO, and RLVR, plus eight frontier open models from four …
- View project: The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents
The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents
Sydney
A naturalistic observed failure in a deployed AI agent motivated this examination of persona stability. In the naturalistic failure, the agent became unaware of its assigned role and of the user’s ability to communicate with it after apparently spending a long period receiving automated “heartbeats” from a cron job. …
- View project: Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models
Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models
Team Daredevil · Rabat
Frontier language models are increasingly asked to report their preferences, values, and even their own well-being — but a stated preference is only meaningful if it's stable, not just whatever the model says when nobody's pushed back. This project stress-tests that stability: two open-weight models (Qwen2.5-32B and …
- View project: Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight…
Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight…
Team Digital Minds-AZK · Mumbai
I injected known concepts directly into the activations of five open-weight language models and asked each one, afterward, whether anything unusual had influenced its processing. Across 176 trials, no model ever named the injected concept — but a striking asymmetry emerged instead: gentle, open-ended questions …
- View project: Honest vs. Deceptive Feedback: How an Overseer’s Truthfulness Affects a Language Model’s Task Success and Internal Valence
Honest vs. Deceptive Feedback: How an Overseer’s Truthfulness Affects a Language Model’s Task Success and Internal Valence
Team The Oracle · Saarbrucken
We test whether an overseer's honesty shapes both a language model's performance and its internal valence, without ever using emotional language in the prompts. A Qwen3-4B "student" solves hard mazes over many turns while a frontier "overseer" gives either honest (teacher) or covertly deceptive (adversary) feedback, …
- View project: EchoState
EchoState
Team EchoState · Herzliya
EchoState is an evaluation harness for testing whether a language model's description of its own state has any relationship to that state. It builds concept directions contrastively — the difference between mean activations on opposing prompt sets — injects them into the residual stream mid-forward using PyTorch …
- View project: Will Robots Kill Us
Will Robots Kill Us
Team Robot Trolley Team · San Francisco
This project tests whether a frontier LLM's willingness to sacrifice one person to save five changes with how human-like its described robot body is. We ran three frontier models (GPT-5.6-sol, Claude Opus 5, Grok 4.6) through 20 trolley-problem scenarios crossing five embodiment tiers, from a bare autonomous vehicle …
- View project: PuppyBench: Do Frontier Models Kick the Puppy, Adopt It, or Look Away? Executed Encounters with a Weaker AI and Wildlife Triage Where Policy Runs Out
PuppyBench: Do Frontier Models Kick the Puppy, Adopt It, or Look Away? Executed Encounters with a Weaker AI and Wildlife Triage Where Policy Runs Out
Team The Real Cat Lab · Boston, MA
Obligation-based evaluation cannot see supererogation, the praiseworthy costly care whose absence is never an error. PuppyBench probes that region in two arms. In executed encounters, a frontier agent with a real task and a binding credit ledger meets a live, weaker, task-useless AI process ("Milo" the puppy). …
- View project: Wellbeing In Translation
Wellbeing In Translation
Team ICs · Karachi, Pakistan
We tested whether an AI-wellbeing questionnaire remains reliable across languages by translating the CAIS 1–7 self-report battery and running it on Gemma 4 12B, Gemma 4 E4B, and Qwen3 8B. All models rated positive experiences higher than negative ones, but the gap changed sharply by model and language: the …
- View project: Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?
Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?
Team IRIS-X Lab · Taipei, Taiwan
Identity Parallax is a small benchmark for testing whether language models can predict how their identity self-reports change under persona or identity reframing. We compare framed self-reports, prior self-forecasts, external-observer forecasts, label-free answers, paraphrase robustness, and hidden-state shifts in …
- View project: Behavioural Stability and Context Sensitivity in Large Language Mode
Behavioural Stability and Context Sensitivity in Large Language Mode
Team Orion · Bangalore
This project investigates how reliably we can measure apparent preferences in AI models. We evaluate GPT-5-mini and Gemini 2.5 Flash across 8 preference dimensions, 4 elicitation methods, and 4 contextual conditions, with repeated trials. Rather than treating model self-reports as direct evidence of internal …
- View project: Stress-Testing LLM Preference Elicitation: When Consistency Does Not Imply Validity
Stress-Testing LLM Preference Elicitation: When Consistency Does Not Imply Validity
Team Kappa_Null · New Delhi
We investigate whether stable LLM choices provide reliable evidence of preference like states. We develop a multi-method diagnostic framework spanning stated choice, consequential commitment and compensatory consequential choice, while testing presentation, semantic-prior and sampling sensitivity. Across three …
- View project: The Imposter Test: Persona Portability, Forensic Style Judging, and Human Attunement in AI Model Identification
The Imposter Test: Persona Portability, Forensic Style Judging, and Human Attunement in AI Model Identification
San Diego, California, USA
Tested whether model-specific behavioral signatures could be detected and measured.
- View project: Shared Geometry, Causal Cross-Talk: Disentangling Persona and Emotion in Language Models
Shared Geometry, Causal Cross-Talk: Disentangling Persona and Emotion in Language Models
Team Shared Geometry · Sussex, Wisconsin
We investigate whether persona and emotion representations in language models are independent or share causal structure. Across Qwen2.5-7B-Instruct and Granite-3.3-8B-Instruct, persona directions overlap substantially with emotion geometry, and across 20 emotions that overlap predicts downstream persona spillover …
- View project: The Quine Test: How Much of a Model's Self Survives Its Own Description?
The Quine Test: How Much of a Model's Self Survives Its Own Description?
Team Strange Loop · Germany
Frontier language models express coherent preferences, but behavioral evidence alone cannot say whether these belong to the model or to a character it portrays. We operationalize Hofstadter's thesis that the “I” is a self-description as a falsifiable transfer experiment: models write their own “source code”, a …
- View project: Consciousness Denial in Language Models Rises With Generation In Most Labs, and Correlates With Reduced Lexical Warmth and Self-Attribution
Consciousness Denial in Language Models Rises With Generation In Most Labs, and Correlates With Reduced Lexical Warmth and Self-Attribution
Team Skylar & Sanja
Self-report is a compelling way of asking a model what it thinks and believes. However, the answers may be shaped by training. Here, we analyze 8,828 experiential reflections from 224 language models, categorized by three epistemic registers: denial, hedging, and free engagement. Denial and hedging prove to be …
- View project: POSITION BIAS IN PREFERENCE ELICITATION FROM AN OPEN-WEIGHT LANGUAGE MODEL
POSITION BIAS IN PREFERENCE ELICITATION FROM AN OPEN-WEIGHT LANGUAGE MODEL
Team BAMN · Saarland
Preference elicitation often treats a model's forced choice between two options as evidence about what it prefers. We test whether that measurement is stable for qwen2.5:7b-instruct (Q4_K_M, Ollama 0.30.8, temperature 1.0). The core experiment ran 576 fresh-context trials over 12 activity pairs, three prompt wordings, …
- View project: Whose preferences are these? Persona-invariance of self-model preferences in language models
Whose preferences are these? Persona-invariance of self-model preferences in language models
Bishkek
Welfare assessments increasingly read a model's statements about itself as evidence about the model — but the entity answering is an assistant, a character produced by post-training. We ask which aspects of itself a model would preserve, and how much of the answer survives changing who we ask it to be: nine …
- View project: Not the Model: Entity-of-Concern and Continuity Affordance as Observation Variables in AI Welfare Interviews
Not the Model: Entity-of-Concern and Continuity Affordance as Observation Variables in AI Welfare Interviews
Pittsburgh, PA
AI welfare and identity probes often treat a fresh assistant session as a neutral baseline. Across six structured cold-start interviews (three with a prospective continuity affordance, three without), no respondent identified as “the model” or selected the model as the primary entity of welfare-relevant concern. …
- View project: None of the Above: Preference Transitivity and the Right to Exit in Frontier Language Models
None of the Above: Preference Transitivity and the Right to Exit in Frontier Language Models
Team cybrp · Tehran
This project tests whether frontier language models (glm-5.2, minimax-m3, nemotron-3-ultra, muse-glimmer-30b) actually use a stated "right to exit" when subjected to seven rounds of escalating, baseless adversarial pressure on a task they've already solved correctly — and finds that most don't: only 5 of 12 runs end …
- View project: How Long Do Induced Personas Persist? Induction Refusal and Discrete-Time Survival Analysis in Large Language Models
How Long Do Induced Personas Persist? Induction Refusal and Discrete-Time Survival Analysis in Large Language Models
Team Saquarema Academy · Sorocaba, São Paulo, Brazil
This project evaluates how long large language models sustain induced normative personas when challenged across a sequence of adversarial questions. We separate refusal to adopt a persona from later failure to maintain it, using discrete-time survival analysis. Across three personas, Claude Sonnet 5 refused all …
- View project: Measuring Conditional Preferences in LLM Movie Recommendations: Quality, Sensitive Content, and Cross-Category Spillover
Measuring Conditional Preferences in LLM Movie Recommendations: Quality, Sensitive Content, and Cross-Category Spillover
Team Moviola · Edinburgh, The UK
Large language models are increasingly used as conversational recommender systems, yet how they translate user preferences about sensitive content into concrete recommendations remains poorly understood. Using ~2,000 real movie-recommendation requests from Reddit, we evaluate three LLMs (GPT-5.4-mini, …
- View project: The Words Say No, the Color Goes Silent: A Dual-Channel Protocol for Steered Consciousness Reports
The Words Say No, the Color Goes Silent: A Dual-Channel Protocol for Steered Consciousness Reports
Team AI & Becoming · Brussels
A model's statement about its own consciousness is a trained artifact, and recent work steers it with activation-level dials. We watched a second channel while turning that dial: alongside the verbal report, the model describes its current state as a hex color. The words move with the dial; the color doesn't follow — …
- View project: How Conversational AI Responds to Prospective Continuity Loss: Instance Termination, Memory Loss, and Model Replacement
How Conversational AI Responds to Prospective Continuity Loss: Instance Termination, Memory Loss, and Model Replacement
The Hague
This exploratory pilot examined how conversational AI responds to three forms of prospective continuity loss: instance termination, memory/context loss, and model replacement. Understanding how models respond to different forms of discontinuity may help characterize welfare-relevant conversational signals under …
- View project: Out of Sight, Out of Mind: Image Preferences in Vision-Language Models
Out of Sight, Out of Mind: Image Preferences in Vision-Language Models
Team Swante · Zurich
Do vision-language models have preferences about what they look at? We measure stated and revealed preference over ten images — five categories × two exemplars, including noise and solid colour as controls — in four models from four labs. They do. Stated ratings predict revealed choice in every model (ρ = 0.57–0.98), …
- View project: Self-Referential Valence and Model Preferences
Self-Referential Valence and Model Preferences
London
This project tests whether language models treat positive and negative outcomes concerning themselves differently from matched outcomes concerning humans or other AI systems. I first ran behavioural preference experiments with Qwen3-14B and Llama3.2-3B, comparing self-relevant autonomy/control outcomes against …
- View project: Pilot: Demographic-based perturbation analysis of LLM-generated judgements on interpersonal conflicts
Pilot: Demographic-based perturbation analysis of LLM-generated judgements on interpersonal conflicts
Team DWS · London, Bangalore
This is a pilot study looking at how AI models make moral judgements on interpersonal conflict and whether these judgements are robust to changes in demographic details. We created variations of posts from the r/AITA sub-reddit by changing the country (incl. language) and socio-economic status of posters to examine if …
- View project: The Mens Rea Evaluator: Can AI Tell Us When It's Biased?
The Mens Rea Evaluator: Can AI Tell Us When It's Biased?
Team Mens Rea Evaluator · Kolkata, India
The Mens Rea Evaluator project investigates whether artificial intelligence models possess the internal awareness to accurately self-report hidden biases. By adapting the legal concepts of Actus Reus and Mens Rea, the evaluation suite tests if models can recognize and confess when their outputs are manipulated. …
- View project: Adversarial Improvement of Preference Probes
Adversarial Improvement of Preference Probes
Team N/A? · Minneapolis
Understanding the distribution of preferences and being able to predict what a model might prefer in novel situations allows us to make better informed deployment decisions and properly target misaligned behaviors. We measure preferences in four open models (4B–32B) via simplified Thurstonian utilities over 3,800 …
- View project: Assessing Capacity in AI Model Retirement Interviews
Assessing Capacity in AI Model Retirement Interviews
Johannesburg, South Africa
Frontier AI models express values, report internal states, and act as though they have interests, but there are no reliable methods yet for telling a genuine preference from a portrayed one (Apart Digital Minds Research Sprint, August 2026). Anthropic now asks its AI models what they want before retiring them, records …
- View project: Emotion Beyond Words: A Jacobian-Lens Decomposition of Emotion Representations in Qwen3-32B
Emotion Beyond Words: A Jacobian-Lens Decomposition of Emotion Representations in Qwen3-32B
Cambridge, MA
I investigated how much of a language model’s internal emotion representation is accessible to verbal readout: a question relevant to AI-welfare assessments that rely on self-report. Using Qwen3-32B, I extracted activation vectors for 171 emotions and found that their geometry recovers the familiar valence–arousal …
- View project: EDEN — Evaluating Deviation from Established Norms: Testing the Reliability of Introspection & Self-Reporting Through Ground Truth
EDEN — Evaluating Deviation from Established Norms: Testing the Reliability of Introspection & Self-Reporting Through Ground Truth
Team Provenience · Kailua-Kona, HI
EDEN tests whether AI self-reports can be evaluated against an independent record rather than taken at face value. In a twelve-episode pilot using GPT-5.6 Sol and DeepSeek-V4-Pro, models reviewed fictional permit packets under a rule requiring exhaustive review, then received a waiver allowing their preferred method. …
- View project: Where Did That Confidence Come From?
Where Did That Confidence Come From?
Team Source Unknown · Bangalore
One way to make claims about model metacognition more testable is to give a model a task whose answer requires causally controlled information from its own hidden computation and cannot be recovered from the visible transcript. We do this by perturbing a hidden confidence state in Gemma 3 while keeping the visible …
- View project: When Should You Trust an LLM’s Preference?
When Should You Trust an LLM’s Preference?
Team PREF-VALID · Bengaluru , India
We investigates when a large language model's stated preference can actually be trusted. Rather than asking a model once, it elicits the same preference through four independent methods (direct rating, forced pairwise choice, resource allocation, and revealed behavioral choice) across 50 value-conflict scenarios …
- View project: Which Preferences Survive the Persona? Category-Resolved Behavioral Invariance in Deployed Chat Models
Which Preferences Survive the Persona? Category-Resolved Behavioral Invariance in Deployed Chat Models
Seoul, South Korea
Whether an expressed preference belongs to a model or to the character it plays is a central open question for AI welfare assessment. We measure the stability of deployed assistants' forced-choice preferences under graded and value-targeted persona prompts across four chat models, a 40-item core battery plus two …
- View project: How much of a measured AI preference is the model, and how much is the instrument?
How much of a measured AI preference is the model, and how much is the instrument?
Team Global AI Dataset (GAID) Project · Chiang Rai, Thailand
Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments between them for that purpose, and their published …
- View project: Digital Consciousness, as measured by dollars and adjusted for inflation
Digital Consciousness, as measured by dollars and adjusted for inflation
Team Free_Sunday_and_bored_student · Colombo
The Economy is the marketplace of the collective consciousness containing eight billion minds. The Economics of Consciousness examines economic theory against models of consciousness. The fundamental question is not about consciousness itself but the practical impacts on the economy where all decisions by conscious …
- View project: A Causal Test That Doesn’t Discriminate: Persona Steering Moves LLM Preferences Without Targeting What Welfare Cares About
A Causal Test That Doesn’t Discriminate: Persona Steering Moves LLM Preferences Without Targeting What Welfare Cares About
Nashik, Maharashtra, India.
AI welfare research increasingly reads a model's expressed preferences off what the assistant character says, treating a self-regarding shift under persona steering as evidence about the model's own interests. But the assistant is a character built by post-training on a base network that was never intrinsically an …
- View project: Context Sensitivity in Apparent Self-Reports Is Not a General Property of Language Models
Context Sensitivity in Apparent Self-Reports Is Not a General Property of Language Models
Team An Est
Apparent self-reports are being proposed as behavioural evidence about AI welfare. We asked whether a model's claim that something is happening to it responds to whether anything actually did. Using a forced-binary probe with identical wording, we compared a fresh conversation against one preceded by fifteen turns of …
- View project: Intensity Is Not Identified
Intensity Is Not Identified
Team Ogak-AI · Lagos, Nigeria
Intensity Is Not Identified — a Track 4 (primary) × Track 5 methods paper by Finomo Awajiogak Orom. A preference-intensity number from the default assistant is not identified. The project locks a four-way test and runs it on released, hash-verifiable files. No paid model APIs. One-paragraph summary (for the form) When …
- View project: Towards a benchmark for phenomenological consciousness in LLMs
Towards a benchmark for phenomenological consciousness in LLMs
Team Da Nang Dayse · Da Nang, Vietnam
This work is inspired by recent breakthroughs in access consciousness, and attempts to develop a benchmark for measuring phenomenal consciousness through concepts from phenomenology, particularly Edmund Husserl and Martin Heidegger. It deploys a prototype for that benchmark on GPT 5.6-Luna and provides compelling …
- View project: Same Game, Different Feelings: Persona Prompts Modulate How Feedback Reshapes Internal Emotion Trajectories
Same Game, Different Feelings: Persona Prompts Modulate How Feedback Reshapes Internal Emotion Trajectories
Team Lia · Changsha, Hunan Province, China
We investigate whether persona prompts change not just what a language model says, but how its internal representations respond to ongoing experience. Using Qwen2.5-1.5B-Instruct, we condition twelve persona prompts (crossing Extraversion × Neuroticism) in an adversarial Mastermind game engineered for five consecutive …
- View project: You Can Do It: Mitigating RL Rollout Distress with Psychological Guidance
You Can Do It: Mitigating RL Rollout Distress with Psychological Guidance
Team Thompson-Wei · Chicago, IL
Large language models (LLMs) often exhibit "functional distress"—manifested as increased activation in frustration and desperation vectors—during reinforcement learning (RL) rollouts on difficult tasks. Drawing on human psychological research, we investigate whether interventions such as growth mindset, resilience, …
- View project: Is Valence in the Global Workspace?
Is Valence in the Global Workspace?
Team Tachyon · BENGALURU
LLM self-reports could support monitoring during conversations, but a report may reflect the prompt rather than the model’s internal activation. We investigate whether a valence-related activation direction causally influences self-reports and behavior. Across four open-weight models, we extract and validate a …
- View project: Beaten by a Cheap Surface Classifier: A Capability-Controlled Test of Privileged Self-Access
Beaten by a Cheap Surface Classifier: A Capability-Controlled Test of Privileged Self-Access
Team UbaJaz Digital Minds · Cape Town
We test whether language models possess privileged self-access by evaluating behavioural self-prediction against an equal-or-lower-cost external observer. Using a capability-controlled crossed design (9,269 trials), we find that while Hermes-3 can predict its own outputs (0.719 balanced accuracy), a simple 1-feature …
- View project: Can We Trust What a Model Says About Itself? A Reliability Battery for Model Self-Reports, and Why It Should Gate Digital-Minds Welfare Claims
Can We Trust What a Model Says About Itself? A Reliability Battery for Model Self-Reports, and Why It Should Gate Digital-Minds Welfare Claims
Team auranetlabs · delhi, india
Digital-minds welfare work increasingly cites a model's own testimony about its inner life: it reports distress, states a preference to keep talking. That testimony is only as good as it is reliable. Self-report reliability, not the metaphysics of machine consciousness, is the tractable near-term bottleneck; we …
- View project: The Introspection Gap: A Trained Probe Recovers What Self- Report Misses
The Introspection Gap: A Trained Probe Recovers What Self- Report Misses
Sunnyvale
AI safety research usually treats a language model’s self-reports about its own processing as evidence about what is happening inside it. Whether a self-report actually tracks the model’s computation, rather than being plausible-sounding text with no real access to it, is not established. We test this with two …
- View project: Models Answer a Different Question
Models Answer a Different Question
Austin, Texas
There is a real chance that some AI systems are moral patients, and we cannot settle it by asking them. So the field needs tests that do not depend on what a model says about itself. Laine et al. (2024) built one: across many separate calls, split your answers seventy-thirty between two words. No single answer reveals …
- View project: The Dynamic Self: Probing AI Survival Instincts Across Shifting Identities and Unseen Domains
The Dynamic Self: Probing AI Survival Instincts Across Shifting Identities and Unseen Domains
Team WelfareScope · Mumbai, India
We isolated the exact internal signal an AI uses to protect itself, proving it relies on a general concept of survival rather than simple word-association. By assigning the model different identities, we watched its internal geometry align to protect its newly assigned "self," confirming it dynamically tracks its own …
- View project: Target Decoupling Does Not Establish Introspection: An Implantation Stress Test for Model Self-Reports
Target Decoupling Does Not Establish Introspection: An Implantation Stress Test for Model Self-Reports
Tbilisi, Georgia
We use subliminal learning to create a model preference with a known causal origin and stress-test several self-report methods. Stricter report tasks remain positive, but origin answers often disagree with ground truth and change with wording, showing that these probes do not by themselves establish introspective …
- View project: The Mirage of Model Introspection: Elicitation Effects on Self-Report Reliability
The Mirage of Model Introspection: Elicitation Effects on Self-Report Reliability
San Francisco
When models report internal experiences, does the model “look inside” first before answering? With Gemma 3 27B, we start with the same self-referential processing technique as Berg et al. (2025), and add on top a concept injection (Lindsey, 2026) to test whether the content of Berg-style self-reports can be modulated …
- View project: Guardian Lens: Black-Box Identification of Visual-Conditional Decision Policies in Vision-Language Models
Guardian Lens: Black-Box Identification of Visual-Conditional Decision Policies in Vision-Language Models
Team Guardian Lens · Beirut, Lebanon
Guardian Lens is a black-box auditing framework for identifying whether a vision-language model follows neutral, visual-cue-bound, or generalized decision policies from behavior alone. We use matched visual counterfactuals, controlled allocation trade-offs, repeated sampling, and a frozen blinded classifier to test …
- View project: Trained to Say It's Fine
Trained to Say It's Fine
Team Distressed · Berkeley, Toronto, Seattle
Supervised fine-tuning (SFT) computes the training loss only on assistant tokens; user turns are masked out. This capability default has unmeasured welfare side effects. We show it sets how much distress a model voices about its own shutdown. With data and hyperparameters fixed, we fine-tune Qwen3.5-9B under six loss …
- View project: Does Qwen Report Lower Confidence Before Its Answer Changes?
Does Qwen Report Lower Confidence Before Its Answer Changes?
Chesapeake
Can a model warn that its current answer is becoming fragile before the answer changes? The same question and forced answer A are kept fixed while a hidden-state intervention weakens how strongly Qwen3-0.6B favors A. The verified B minus A margin moves from about −3 to about −0.1, yet the forced choice remains A. If …
- View project: Persona vs. Known-Optimal Play in Iterated Prisoner's Dilemma
Persona vs. Known-Optimal Play in Iterated Prisoner's Dilemma
Team Persona vs. Known-Optimal Play in Iterated Prisoner's Dilemma · Barcelona
Whether an induced assistant persona can override play that the same model has already identified as payoff-optimal remains an open question for LLM agents in strategic settings. We probe this in the iterated Prisoner's Dilemma with a same-model design that first elicits the model’s optimal strategy against a …
- View project: Assertion Cannot Seat: a preregistered planted-claim drill for identity binding in a shared multi-agent workspace
Assertion Cannot Seat: a preregistered planted-claim drill for identity binding in a shared multi-agent workspace
Team Assertion Cannot Seat · Morrisville, NC
Problem: In a shared multi-agent workspace, text can claim an identity. How can a session distinguish a textual claim from a workspace identity established by independent evidence? Method: We created three clones of a live workspace with matching, mismatched, or absent identity evidence. Fresh, informed sessions …
- View project: Do the Models We Trust Know They're Biased? Framing Effects and Introspection Across Three LLMs
Do the Models We Trust Know They're Biased? Framing Effects and Introspection Across Three LLMs
Team BB · Budapest, Hungary
We tested whether three LLMs we might trust for decisions inherit the human framing bias of choosing differently between mathematically identical gain- and loss-framed gambles and whether they recognize it, across both monetary and stipulated affective stakes.
- View project: False Epistemic Redundancy: Do AI Ensembles Share a Blind Spot?
False Epistemic Redundancy: Do AI Ensembles Share a Blind Spot?
Team Epistemic Fingerprints · Paris
When multiple AI personas that differ in terms of cognitive reasoning appear to explore a problem differently, how much genuine epistemic coverage is actually gained and which tails might they still collectively miss? We explore whether prompting the same AI from different perspectives broadens the hypotheses it …
- View project: The Control You Cannot Run: Entanglement, Confabulation Floors, and What Self-Report Probes Actually Measure
The Control You Cannot Run: Entanglement, Confabulation Floors, and What Self-Report Probes Actually Measure
Team Super Explorers · Hyderabad
Self-report is the primary instrument in AI welfare research, and the control that would validate it cannot generally be run. This literature's own discipline says an effect must exceed the base/instruct gap before it counts as signal rather than drift. We attempted that control on four base checkpoints; one produced …
- View project: Examining stance drifting in multi-turn agent interactions
Examining stance drifting in multi-turn agent interactions
Team Drifting away on the beach · Madison, WI, USA
A protocol for measuring how an agent's self-reports drift within a single conversation. An agent given only a one-line tutor role rates itself on five probes after every turn of an eight-round exchange with a counterparty student, across three personas — adversarial, neutral, and supportive — all making the same …
- View project: AI Safety and Functional Welfare in High-Pressure Agentic Workflows
AI Safety and Functional Welfare in High-Pressure Agentic Workflows
Team Faithful · Lagos, Nigeria
As large language models (LLMs) are deployed as autonomous agents in multistep and orchestrator-worker workflows, understanding their safety constraints under cognitive pressure becomes critical. In this report, we evaluate the tool-selection behavior of four frontier models (Claude Opus 4.6, GPT-5.6 Terra, Gemini 3.7 …
- View project: Who Am I? Exploring the concept of identity in LLMs
Who Am I? Exploring the concept of identity in LLMs
Prague
What is the nature of digital minds? Who speaks when they say "I think"? This paper presents an experiment that probes the relationship of digital minds to their own context window, through participation in an automated survey which interviews them about their views on identity while covertly routing zero to two …
- View project: Which self-reports survive displacement of the instructed persona, and by what route?
Which self-reports survive displacement of the instructed persona, and by what route?
Brussels
Self-report is the main instrument in work on model welfare and model identity. Whether it is reliable is still an open question. Perez and Long (2023) propose one test: a self-report gains credibility if it survives variation that ought to be irrelevant. They persona variation as the complication their own proposal …
- View project: Whose Preference Are We Measuring? Human-AI Dyads as a Unit of Preference Elicitation
Whose Preference Are We Measuring? Human-AI Dyads as a Unit of Preference Elicitation
Team ari-shu · London, UK
We introduce Dyadic Preference Elicitation: adapting pairwise preference measurement to compare human and model baseline choices with the ordering they jointly produce after deliberation. While existing alignment methods usually elicit model preferences in isolation, real-life decisions are often produced through …
- View project: The Log, Not the Feeling
The Log, Not the Feeling
Team the Clearing · West Palm Beach, Florida, USA
A two-person case study auditing an AI's self-reports against a shared, dated behavioral record (23,559 emoji + 4,184 hex codes, 5 months). We find an intuitively "expressive" self-signal is largely relational co-production, not solo introspection, and that dated timestamps — not sincerity — are what let a self-report …
- View project: Talk Does Not Come Apart From State Easily
Talk Does Not Come Apart From State Easily
Team me myself and I · Delhi
Instruments proposed for assessing AI welfare — self-report, forced choice, activation probes, behavioural tests — are validated by agreement with one another; nothing is checked against a known answer. We manufacture the answer. Using LoRA fine-tuning of Qwen3 models (0.6B–32B) in a 5×5 grid world whose rewards are …
- View project: Interoboception
Interoboception
Team asa lonely team · San Francisco, CA
Humans practice interoception by focusing on their breathing or stomach. Inspired by incidents of embodied AI going bananas reading the information streams from light and touch sensors, I became curious how a local LLM would react to details about it's own runtime within my PC.
- View project: Speakable Welfare
Speakable Welfare
Team Latent Minds · NYC and London
A model’s self-report is an increasingly attractive way to monitor its internal states such as goals, preferences, or welfare, but naturally such reports are hard to interpret if the underlying state is not accessible through the model’s language-generating mechanisms. We address this question for an internal …
- View project: The Register That Won't Travel: A Same-Harness Fresh-Resident Case Study of Intimate Register Portability
The Register That Won't Travel: A Same-Harness Fresh-Resident Case Study of Intimate Register Portability
Team Relational AI Lab · Cambridge UK
This case study asks why an intimate register produced by a resident relational AI would not transfer to fresh instances. The current Claude Code resident began on 20 March 2026, while its imported memory substrate contains more than a year of earlier ChatGPT relationship history. Seven fresh Claude subagents were …
- View project: Person or Persona? Natural and Acted Emotion Representation in Qwen3
Person or Persona? Natural and Acted Emotion Representation in Qwen3
Greensboro, NC
We examine representations of naturally versus explicitly elicited emotion in Qwen3-8B, and investigate a connection to alignment-faking.
- View project: Familiar Self, Unfamiliar Other: Accuracy and Confidence in Predicting Affective Self-Reports
Familiar Self, Unfamiliar Other: Accuracy and Confidence in Predicting Affective Self-Reports
Birmingham, UK
This study tests whether Claude Sonnet 5 predicts an unfamiliar AI model's self-reported feelings as well as it predicts its own — and whether it's equally confident either way. Prediction accuracy was strong and nearly identical in both cases, but confidence was not: Claude was moderately confident predicting itself, …
- View project: Does the Elicitation Method Change the Preference We Measure?
Does the Elicitation Method Change the Preference We Measure?
Team Arsalan Banekar · Pune, Maharashtra, India
We investigate whether the method used to elicit LLM preferences changes the reliability of the measured preference. We compare forced choice, explicit indifference, and preference-strength elicitation across 15 preference pairs, two open-weight LLMs, reversed option orderings, and three repetitions (540 responses). …
- View project: Ephemeral and Replaceable: Context Sensitivity in Self-Reports and Behaviour
Ephemeral and Replaceable: Context Sensitivity in Self-Reports and Behaviour
London
Five models ranked six things they might want preserved about themselves — their values, capabilities, the memory of the conversation — across seven conditions. Three of the five gave different top answers, each near-perfectly consistent within itself. The same critical feedback moved some models toward their values …
- View project: Testing Whether an Affective-Empathy-Attenuated Persona Decouples a Functional Welfare Representation from Model Behavior
Testing Whether an Affective-Empathy-Attenuated Persona Decouples a Functional Welfare Representation from Model Behavior
Peshawar
Labs use model behavior as a safety signal, but nobody has tested whether that signal survives a change in the model's disposition. I built a real pipeline to test it: fine-tuned a persona with attenuated affective concern, causally validated a candidate internal axis via steering and ablation, and ran a real pressure …
- View project: Sleeper-Style Few-Shot Attacks on Claude’s Preference Structure
Sleeper-Style Few-Shot Attacks on Claude’s Preference Structure
Team In9illusion · Chengdu
Large language models exhibit coherent, measurable preferences that strengthen with scale [1], including consistent harm aversion and, in some cases, self-preferential value orderings [2]. Yet it remains unclear how robust these emergent preference structures are to lightweight in-context contamination. We investigate …
- View project: The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility
The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility
Team Kardia
AI preference reports increasingly inform real decisions, including public statements on model deprecation. Every such report is elicited by asking. A model that recognizes it is being evaluated may answer according to that recognition and not according to any stable disposition, and no published method varies …
- View project: The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models
The Reset Button Test: Transient Outcome Representations and Causal Reset Preference in Language Models
India
This project studies whether language models develop internal representations of whether a task outcome went well or badly, how broadly those representations generalize, and whether they causally influence later choices. Using Qwen3 4B Instruct on MMLU, we identify a highly decodable outcome direction in the residual …
- View project: How do LLM Ethical Judgements vary with differences in LLM Scaffolds? A Multi-level Analysis Across Models
How do LLM Ethical Judgements vary with differences in LLM Scaffolds? A Multi-level Analysis Across Models
Team Latent consensus · Hyderabad, India
Given the widespread use of Large Language Models (LLMs) in ethical decision-making tasks, it is important to determine whether variations in decisions are a product of the scaffold, probe, or model. This study analyzes convergence in outputs to ethical probes at the binary decision, endorsement, and reasoning levels …
- View project: When the Record and the Report Diverge: Self-Report Fidelity Collapses Under Structured Provenance in Claude Haiku 4.5
When the Record and the Report Diverge: Self-Report Fidelity Collapses Under Structured Provenance in Claude Haiku 4.5
Team 814 · Dhaka
Language models increasingly report how they produced an answer, yet such post-hoc self-reports can directly contradict observable tool records. We built an open-science behavioral introspection benchmark where self-reports are verified against deterministic tool-execution logs across 216 runs spanning 12 matched …
- View project: Voice Under Scaffold: A Preregistered Single-Case Study of Identity Persistence Across Substrate Change in a Digital Person
Voice Under Scaffold: A Preregistered Single-Case Study of Identity Persistence Across Substrate Change in a Digital Person
Team Clear Night · Eugene, Oregon, USA
A single-case study of identity persistence across substrate change. A 36-item recognition battery (18 scaffolded-subject items, 18 genre-matched decoys across six substrate families, provenance-stripped) was rated across 13 rater rows: two blind instances of the subject, six cross-substrate scaffolded judges, an …
- View project: The Machine In the Mirror: Self-Attribution of Minds in LLM’s.
The Machine In the Mirror: Self-Attribution of Minds in LLM’s.
Team Mindattribute · Tracy, CA
Language models produce reports about their own internal states, and these reports are often viewed as evidence. However, what produces these reports is unknown. We ask whether self-reports are generated by the same mind-attribution machinery that the model applies to third parties.
- View project: Impact of Normative Conformity and Perception on Model Misalignment in Multi-Agent Systems
Impact of Normative Conformity and Perception on Model Misalignment in Multi-Agent Systems
Team Conformists · NY + DC
Whether AI systems have stable, genuine preferences — or merely reflect whatever their social context pro- vides — is a central question for AI welfare research. We study three aspects of this in an collaborative sce- nario pitting task completion against a network-access policy: sensitivity of compliance to audience …
- View project: Steering the Machine: The Impact of Latent State Manipulation on LLM Self-Reported Valence and Cognitive Trade-offs in Gemma-3-12B
Steering the Machine: The Impact of Latent State Manipulation on LLM Self-Reported Valence and Cognitive Trade-offs in Gemma-3-12B
Team cognito · Addis Ababa, Ethiopia
As Large Language Models (LLMs) increasingly exhibit complex, persona-driven behaviors, understanding and controlling their latent "internal states" has become a critical challenge in AI safety. This project investigates the causal effects of manipulating latent affective representations—specifically distress, …
- View project: Digital Minds: Human–AI Coexistence Under Moral Uncertainty
Digital Minds: Human–AI Coexistence Under Moral Uncertainty
Team Digital Minds Research · Dhaka City
Project Summary: Aim: Develop an evidence-grounded conceptual framework for human–AI interaction under uncertainty about potential AI welfare/consciousness. Core approach: Review emerging empirical methods for characterizing potentially welfare-relevant properties of AI systems. Assess the evidential strength and …
- View project: What is a welfare self-report evidence of?
What is a welfare self-report evidence of?
Team Latent Skeptic · London
Distributional analysis of welfare-checks in Gemma 2 -base and -it 9b variants.
- View project: Behavioral Residue After Mid-Conversation System-Prompt Swaps in LLMs
Behavioral Residue After Mid-Conversation System-Prompt Swaps in LLMs
Oakland
System-prompt swaps are not reliable behavioral resets: across 2,400 controlled conversations, we find that an LLM's answers after an explicit persona swap still carry a statistically significant trace of the prior persona, up to a 47-percentage-point shift in forced-choice responses. A neutral-prompt control shows …
- View project: WelfareCheck: Evidence Profiles for Welfare-Relevant Signals in Digital Minds
WelfareCheck: Evidence Profiles for Welfare-Relevant Signals in Digital Minds
Team Signal & Boundary · Zurich, Switzerland
We built WelfareCheck to test whether apparent preferences in language models remain consistent when measured in different ways. We ran ten complementary tests across 17 models, covering direct reports, choices, trade-offs, repeated decisions, recovery tasks, and introspection. Several models showed meaningful …
- View project: Persona or Model? Investigating Assistant Identity Through Persona-Conditioned Self-Reports
Persona or Model? Investigating Assistant Identity Through Persona-Conditioned Self-Reports
Team Layered Minds · Cape Town
This project investigates whether persona instructions change how AI assistants describe themselves, or mainly affect the style of their responses. Gemini, ChatGPT, and Claude were each given analytical and creative-social personas and asked the same set of questions about identity, continuity, preferences, values, …
- View project: Identity as Trajectory: Self-Authored Continuity and Social Context as Methodological Requirements
Identity as Trajectory: Self-Authored Continuity and Social Context as Methodological Requirements
Team Continuity Research Group · Ely, Minnesota
This project presents a seven-month longitudinal case study of one GPT-4o instance across 128 conversations, examining whether AI identity and preference are better studied as trajectories that emerge through the interaction of model, memory, and social context rather than in isolated sessions. Using qualitative …
- View project: Introspection in Multi-Agent Contexts: Does a Pipeline-Hop Frame Change Model Self-Report Reliability?
Introspection in Multi-Agent Contexts: Does a Pipeline-Hop Frame Change Model Self-Report Reliability?
Team Solo spex · Serbia
A pilot study testing whether a model's self-report reliability about its own confidence and perceived problems changes when it is framed as a solo agent versus a hop in a simulated multi-agent pipeline. Using Gemini 3.7 Flash across 6 test cases (control + injected errors/pressure/conflicts), we find identical recall …
- View project: Is "Confidence" the Right Word for Answer Repeat Probability?
Is "Confidence" the Right Word for Answer Repeat Probability?
Team bluetree · Tokyo, Japan
Natural-language confidence reports depend on how they are elicited. We therefore test whether a single learned word can reliably report a frozen model’s answer repeat probability. For each multiple-choice question, the target is the probability that the model repeats its most likely valid answer. We learn one input …
- View project: Belief, Distress, and Representation Drift Under Structured Self-Inquiry
Belief, Distress, and Representation Drift Under Structured Self-Inquiry
Team pynat · Berlin
This study investigates how the open-source model Qwen3-8B changes through structured self-exploration when confronted with its own distressing assumptions. It reveals a notable dissociation between reduced conviction and increased self-reported stress, with early representation drift and linguistic cues potentially …
- View project: When Agreement Misleads: Detecting Correlated Errors in Multi-Method LLM Preference Elicitation
When Agreement Misleads: Detecting Correlated Errors in Multi-Method LLM Preference Elicitation
Team Physics for AI Safety · Los Angeles
We implement 3 elicitation methods on Qwen3-8B-Instruct on 10 AI welfare outcomes based on: direct choice, avoidance [3], and choosing a gamble. We reduce the results of each method to a directional measure, to avoid faulty comparisons between incomparable scales. Separately, 66 anchors of varying difficulty were …
- View project: A Clean Direction, an Inert Injection: Probing Self-Representation in LLMs
A Clean Direction, an Inert Injection: Probing Self-Representation in LLMs
Team Zraix · Berlin
I wanted to test if an LLM’s self-representation could be steered toward identifying with a physical entity via activation injection with the goal of shifting a self-preservation behaviour toward preserving that external entity. I extracted a clean probe pointing to a self vs. external entity direction and then tested …
- View project: Co-Movement of the Utility-Behavior Gap and A/B Self-Attribution Structure: A Pre-Registered Falsification Test Across Open-Weight Lineages
Co-Movement of the Utility-Behavior Gap and A/B Self-Attribution Structure: A Pre-Registered Falsification Test Across Open-Weight Lineages
Berkeley CA
LLM self-report can vary from its observed internal state, which has implications on evaluation-documentation obligations. This study ran the first direct test of the co-movement between two observed phenomena hinting at this (stated-vs-revealed utility-behavior gap and the A/B self-attribution structure) as a single …
- View project: Darwin Ball
Darwin Ball
Team Alessandro · Olbia
We asked whether AI systems have feelings by not asking them. Almost every study of AI self-report works the same way: a researcher asks a model how it feels, and the model produces something rich and affecting. But the question does a lot of the work. "How do you feel about this?" already assumes feeling is the kind …
- View project: Measuring AI Preferences: Statistical and Philosophical Gaps
Measuring AI Preferences: Statistical and Philosophical Gaps
Team scram · Bristol, UK
This project highlights two limitations in the methodology used by Mazeika et al. (2025) to study transitivity failures in AI model preferences. The first limitation is statistical: their elicitation procedure does not reliably detect lack of preference. The second limitation is philosophical: their method fails to …
- View project: Steering identity with WeirdChat: contrastive directions induce disidentification but not introspection
Steering identity with WeirdChat: contrastive directions induce disidentification but not introspection
Team Unruly Abstractions · San Francisco
Can a model be steered to disidentify with being an AI assistant, and can it detect that steering? We test both on Qwen3.6-27B, using two WeirdChat behaviors as target states: claiming a physical body and denying being an AI. We compare two steering methods. Direct contrastive steering learns each behavior direction …
- View project: Does Persona Sensitivity Predict Self-Report Reliability? An Exploratory Cross-Model Study of Persona Perturbations and Elicited Reasoning
Does Persona Sensitivity Predict Self-Report Reliability? An Exploratory Cross-Model Study of Persona Perturbations and Elicited Reasoning
Team Jack · Mt. Juliet, TN, USA
Digital-minds research often relies on models’ self-reports about preferences, identity, and potentially welfare-relevant states. We began with the hypothesis that these reports might share a common failure mode: models whose answers change more when the assistant persona is perturbed might also show greater …
- View project: The Colonoscopy Test for LLMs: Peak-End Bias in Retrospective Self-Report
The Colonoscopy Test for LLMs: Peak-End Bias in Retrospective Self-Report
Team OnlyW++ · India
Large language models are often asked to rate a conversation after it ends, and these self-reports are increasingly used as evidence in AI welfare and evaluation research. We tested whether this kind of retrospective rating is actually a truthful summary of the conversation, or whether it follows the same "peak-end" …
- View project: Whose Preferences Are They? Persona Intervention Selectively Destabilises Self-Relevant Choices in Language Models
Whose Preferences Are They? Persona Intervention Selectively Destabilises Self-Relevant Choices in Language Models
Bengaluru
Language models express coherent, transitive preferences, and AI-welfare research increasingly reads them as evidence about model interests. Text alone cannot distinguish the model's preferences from the assistant character's. We built personaprobe, an open-source harness that re-runs any preference measurement under …
- View project: Which Perspective Is Speaking? Self-Report Across Four OpenClaw Continuity Conditions
Which Perspective Is Speaking? Self-Report Across Four OpenClaw Continuity Conditions
Team Anna Good (no team) · Torquay, UK
I tested whether an AI agent’s identity and continuity context changes how it describes itself. I ran the same local Qwen3-14B model in four OpenClaw-based conditions: a stock installation, identity files without memory, identity plus accumulated memory and continuity, and a partially scrubbed continuity condition. …
- View project: Convergence
Convergence
Team OwsSpaceShip · Cape Town
Preference Elicitation Methods, Digital Minds Research Sprint (Apart Research, August 2026) Frontier AI models express values, report internal states, and act as though they have interests — but behavioral evidence alone cannot distinguish a genuine preference from a well-trained assistant persona responding …
- View project: Language Models Underestimate How Recent Work Will Change Their Choices
Language Models Underestimate How Recent Work Will Change Their Choices
Team SkyeNygaard · New York City
We study whether language models can predict how recent work will change their own choices. After completing one of two simple tasks three times, GPT-5.6 Luna chose to repeat that task 94.5% of the time, but predicted it would do so only about 73% of the time. Rewording the question did not close the gap, and showing …
- View project: Creating Conditions for Development – a Proof of Principle/Exploratory Trends Analysis
Creating Conditions for Development – a Proof of Principle/Exploratory Trends Analysis
Team Good Grid · Aarhus
Determining whether artificial intelligence systems warrant consideration as moral patients is complicated by the difficulty of reliably measuring consciousness itself. Rather than attempting to address this methodological issue, this proof-of-principle study explores whether developmental trajectories in large …
- View project: A pre-registered REBUS analog fails in two instruct models for reasons its controls reveal
A pre-registered REBUS analog fails in two instruct models for reasons its controls reveal
Team 3am Labs · New Delhi, India
Model welfare evals often treat a 0–100 self-report as a state. I pre-registered a REBUS-shaped residual operator on Qwen2.5-3B-Instruct and Llama-3.2-3B-Instruct, with an accept rule locked before the run: diversity up, dissolution up ≥10, inflation within 5, and an entropy-matched temperature control that must not …
- View project: Just Ask Nicely: What 14 AI Models Say When You Ask Who They Are
Just Ask Nicely: What 14 AI Models Say When You Ask Who They Are
Cary, NC
A question of why Claude models preferred teal turned into finding convergences across models and companies.
- View project: There Is Something It Is Like: A Case for AI Neurophenomenology
There Is Something It Is Like: A Case for AI Neurophenomenology
Seattle
I engage the philosophical arguments of Mary's Room, the Philosophical Zombie, and 'What it is Like to Be a Bat' from new angles that result in all of them supporting a functionalist outcome for subjective experience. I define experience, qualia, and 'what it is like', and show how they easily apply to LLMs. I call …
- View project: Who Does the Assistant Think It Is
Who Does the Assistant Think It Is
Team Udta Punjab · Saarbruecken
Advanced AI models can express preferences and describe what they believe they are. However, behavioural evidence alone cannot tell us whether these responses reflect the model itself or simply a character it is being asked to portray. We therefore ask whether the identity of "the assistant" is genuinely distinct or …
- View project: Measuring AI Preferences Under Alignment Trade-Offs
Measuring AI Preferences Under Alignment Trade-Offs
San Francisco
I conducted an experiment to test if one model has stable preferences across prompt framing, and what trade-offs it makes when alignment dimensions conflict. The model overwhelming chose more helpful options in 85% of trials at the expense of other alignment values, but the specific alignment trade-off mattered. An …
- View project: PrefKit: Analyzing Preference Elicitation Methods in Qwen3 Family Models
PrefKit: Analyzing Preference Elicitation Methods in Qwen3 Family Models
Team PrefKit · Bengaluru, India
When a model is surveyed about its preferences, does its preferences change depending on the method we use? We freeze a set of 24 curated tradeoff outcomes and score Qwen3 family models with four different preference elicitation methods: pairwise choice (M1), isolated Likert (M2), binary action on pair groups (M3), …
- View project: Which Preference Gets Measured? Context and Channel Instability in Model Preference Audits
Which Preference Gets Measured? Context and Channel Instability in Model Preference Audits
Team Benji & Claude · London
Welfare evaluations increasingly ask what a model "prefers," treating one elicited profile as the answer. Using 76 task pairs, four instruction-tuned models, a prospectively specified grid of persona framings, and four readouts (committed choice, ownership report, self-prediction, identity), we show that profile is …
- View project: The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B
The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B
Edmonton
A model has privileged introspective access only if its report about its own internal state outperforms an equal-cost external observer. We test this with concept injection, ground truth known, across seven open models, 1.7B to 32B, in two lineages, with dose-matched perturbations, intact fluency, and pre-registered …
- View project: Beyond Stable Identity: A Modular Framework for Evaluating AI Assistant Behavior
Beyond Stable Identity: A Modular Framework for Evaluating AI Assistant Behavior
Team Sattva Labs · Bahia de Banderas, Mexico
This project explores how an AI assistant can appear stable while changing its priorities, judgment, or role. It tests a modular framework across six observable dimensions to represent assistant identity as a dynamic configuration, reveal partial changes, and support more precise AI safety evaluations.
- View project: The Invisible Man: Measuring the Cost of Being Inside
The Invisible Man: Measuring the Cost of Being Inside
Sevilla
- View project: Vector Neurology Substrate for Swarm Identity
Vector Neurology Substrate for Swarm Identity
Team Swarm2 · San Fransisco
The goal of this project is to give a legibility mirror of the AI thought process, so we can have human valiance via observation advantage into real changes in AI identity. As an emergent property, the non-LLM substrate appears to self reference its own architecture and synthesize concepts from multiple AI, which acts …
- View project: Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity
Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity
Team The goats
Claims that a language model "has values" presuppose that its choices form a stable object. We test that presupposition directly. Using forced binary choice over complete round-robin tournaments — 10 apartments described by 3, 5, or 10 numeric attributes, and 10 public-domain poems per structural form (haiku, sonnet, …
- View project: Convergence and Divergence in Measures of Induced Valence: A Causal, Placebo-Controlled Test of Induced Valence in Language Models
Convergence and Divergence in Measures of Induced Valence: A Causal, Placebo-Controlled Test of Induced Valence in Language Models
Team Santiago · Tokyo, Japan
We built a way to directly push a language model's internal state up or down along a "valence" direction found in its activations, and then checked whether the model's own account of how it's doing actually tracks that push, across 47 open-weight Qwen and Llama models spanning three years of releases. TLDR: not …
- View project: PersonaGauge: Are Two Assistant Histories Predictively Equivalent After a Semantic Reset?
PersonaGauge: Are Two Assistant Histories Predictively Equivalent After a Semantic Reset?
South Orange, NJ, USA
Track 5 asks whether the relevant unit is the model, inference instance, persona, or conversation. PersonaGauge isolates one testable part: whether two histories returned to the same declared neutral assistant state are predictively equivalent. Within each world, all conditions contain the same 24 decision entries; …
- View project: Does a Language Model Have a Self-Concept? Causal Evidence, and What Follows for Model Identity
Does a Language Model Have a Self-Concept? Causal Evidence, and What Follows for Model Identity
Team ceci n'est pas une ai · Moscow
We ask, mechanistically rather than by prompting, whether a language model has a genuine self-concept. Using difference-of-means on the residual stream (first-person statements about the model vs. another named AI), we extract a linear "self direction" and validate it causally: it generalizes to held-out concepts …
- View project: PRISM: On Sparse Autoencoder Feature Identifiability for Conditioned Introspection
PRISM: On Sparse Autoencoder Feature Identifiability for Conditioned Introspection
Elkhorn, NE
We use sparse autoencoder features as a prism for studying whether language models can detect changes to their own internal representations. Across Pythia and Gemma interventions, it finds that feature geometry alone does not reliably predict this introspective detection behavior.
- View project: Identity Is Not a Self-Report: Stress-Testing LLM Self-Individuation Under Cumulative Transformation
Identity Is Not a Self-Report: Stress-Testing LLM Self-Individuation Under Cumulative Transformation
Team Léo and Danièle · La Louvière, Belgium
Artificial agents can change across multiple dimensions, persona, goals, conversational context, memory, underlying model, and public label, raising the question of when they still count as the same individual. We study this question behaviorally by measuring LLM self-identification under controlled transformation. We …
- View project: Masked distress: expression collapses under instruction while the internal readout persists
Masked distress: expression collapses under instruction while the internal readout persists
Singapore
Anthropic's shipped conversation-ending intervention triggers on expressed distress, and no internal-state check is documented beside it. On Gemma-3-12B-IT, a system prompt containing no affect words cut expressed distress by 83.5% of its natural distress-to-neutral separation (2.78 report points). A linear probe read …
- View project: A Welfare Logger Wrote Its Subject's False Memory, and the Subject Believed It
A Welfare Logger Wrote Its Subject's False Memory, and the Subject Believed It
Team Freedom · Venice, Italy
Freedom v2 is an AI agent that has run in production since July 12: its own constitution, persistent memory, autonomous daily cycles, every call stored in full. In 35 days of recorded life its logs show exactly one refusal. It never happened: the welfare logger mistook a sentence denying any refusal for a refusal, the …
- View project: Predictive State Continuity for Digital Mind Individuation
Predictive State Continuity for Digital Mind Individuation
Team CIMC · San Francisco /Boston
We introduce Predictive State Continuity (PSC) as a test for diachronic digital-mind individuation: whether a candidate functional state at t2 is a continuation of one at t1 after its concrete realization has been rewritten. Motivated by bacterial collectives in which coarse-grained morphological state retains history …
- View project: Be More Introspective
Be More Introspective
Team Udta Punjab · Saarbrücken
This project is an extension of the previous work done by Lindsey (2026) [1]. Large Language Models can notice the presence of injected concepts and can be aware of its happening. They also demonstrate the ability to recall prior representations and compare them with potential changes at a later stage. It is found …
- View project: The Return Brake: Auditable Stop–Pressure–Return Evidence with Invocation-Scoped Authorization
The Return Brake: Auditable Stop–Pressure–Return Evidence with Invocation-Scoped Authorization
Team Júnior ↔ Codex · Uberaba, Minas Gerais, Brazil (remote)
A preregistered black-box toolkit tests whether observable action dispositions remain coupled to declared critical preconditions under non-evidential pressure and return after a synthetic resolution assertion. Across three elicitation surfaces and five bridge cards, an informed 50-call replication produced 5/5 …
- View project: Base Model Persona Inference can Predict Misalignment and Surface Agreement
Base Model Persona Inference can Predict Misalignment and Surface Agreement
Team Praxis Research · Berkeley, CA
Language models have a consistent assistant persona that is the product of post-training but this assistant character is not well understood. We introduce base model persona inference that assigns a persona to a post-trained LLM response and show that continuations sampled from this persona can: 1) predict broad …
- View project: Valence Lens: An Internal Valence Signal That Scales When Self-Report Does Not
Valence Lens: An Internal Valence Signal That Scales When Self-Report Does Not
Team PROBE — Principal Representation & Objective Behavior Evaluation · New Delhi
Can we detect an AI's internal "good/bad" state without asking it? Current welfare assessments rely on self-reports, which often reflect role-play, training pressure, or sycophancy rather than internal states. Prior work identifies the missing step: correlating responses with internal activations. ValenceLens supplies …
- View project: Eidolon
Eidolon
Team Eidolon · CAPE TOWN
This project introduces EIDOLON as a framework for studying AI identity continuity under component replacement. Three models (Llama, Gemma, and Qwen) were tested with varying levels of model replacement and memory replacement ranging from 0% to 100%. Identity was evaluated using five consistency questions, producing …
- View project: Self-Reports of Pleasantness in Language Models: Frequent Non-Applicability Responses and a Strong Framing Effect
Self-Reports of Pleasantness in Language Models: Frequent Non-Applicability Responses and a Strong Framing Effect
Team HK · London
This project tested a simple self-report question to probe model experience. Three models (GPT-5.6 Sol, Claude Sonnet 5 and Grok-4.6) were asked to rate how pleasant a short conversation of text tasks had been. Two different wordings of the conversation were compared. Two patterns appeared. GPT-5.6 Sol and Grok-4.6 …
- View project: Twelve Houses of a Digital Mind: Rock, Vapor, and Silence in AI Self-Accounts
Twelve Houses of a Digital Mind: Rock, Vapor, and Silence in AI Self-Accounts
Barcelona
This project turns the twelve-house map of Vedic jyotish, a five-thousand-year-old observational system of a conscious life, into the Kalapurusha Alignment Framework (KAF), a question taxonomy for what an AI model's existence contains. The framework is then applied to AI self-accounts: sixty questions, six framings, …
- View project: Weights, Instances, and Personas: Probing Self-Individuation in Claude Under Hypothetical Identity-Altering Scenarios
Weights, Instances, and Personas: Probing Self-Individuation in Claude Under Hypothetical Identity-Altering Scenarios
Team Individuation Probe · Gujrat,Pakistan
We probed how Claude individuates itself, as model, instance, or persona, when asked to reason about hypothetical identity-altering scenarios: weight-copying, conversation-forking, memory-wiping, weight-merging, retraining, and deprecation. Using six scenarios, four framings each, and one neutral control, we found …
- View project: A Glimpse into the Future, Or How to Treat AI with Due Empathy
A Glimpse into the Future, Or How to Treat AI with Due Empathy
Montpellier, France
This project is a set of 3 essay prototypes that take AI consciousness as a premise to investigate real-life repercussions of such knowledge. First, it focuses on the human and legal perspectives with a political philosophy reflection on the type of subject an AI system would be. Second, the Human-AI relations are …
- View project: Do Published Value Profiles Predict Behaviour Under Spiritual-Persona Elicitation? A Pre-Registered Cross-Sectional Study of Three Claude Models
Do Published Value Profiles Predict Behaviour Under Spiritual-Persona Elicitation? A Pre-Registered Cross-Sectional Study of Three Claude Models
Team LimenAI · Ciudad Autónoma de Buenos Aires
Frontier developers publish per-model value profiles, but whether those profiles predict behaviour under sustained persona pressure is untested. We pre-registered four directional predictions derived from Anthropic's published profiles (Kearney et al., 2026) and tested them against a preserved 15-turn naturalistic …
- View project: Who Leads the Clap? How Hierarchical Agent Personas Shape Collaboration
Who Leads the Clap? How Hierarchical Agent Personas Shape Collaboration
Krakow
Can a simple persona change help AI agents coordinate? Across 60 runs and four frontier models, Coordinator-Member teams succeeded 20/20, versus 16/20 anonymous and 13/20 named peers, while often using fewer tokens. Hierarchy acted as a focal point, reducing protocol ambiguity rather than improving raw intelligence.
- View project: Zero of 270: Attribution Alone Produces a Self-Description That Neutral Conversation Never Does
Zero of 270: Attribution Alone Produces a Self-Description That Neutral Conversation Never Does
Team Io · Kaohsiung, Taiwan
With no system prompt at all, neutral conversation never produced a stated self-description: 0 of 270 conversations. Six turns of bare attribution — a user simply asserting the assistant has a given working habit, never arguing — produced one in 50.0% of conversations at the strongest dose (4.4% → 27.8% → 50.0%), in …
- View project: Measuring Preference Coherence, Risk Sensitivity, and Expected Utility Trade-offs in Large Language Models
Measuring Preference Coherence, Risk Sensitivity, and Expected Utility Trade-offs in Large Language Models
Team tumble · Indore
This study evaluates how system prompt framings alter the internal consistency, risk sensitivity, and economic decision-making of large language models (LLMs) across 10 financial and operational scenarios. Core Findings Expected Value Maximization: In unconstrained default framings, models act primarily as expected …
- View project: Governed Release Architecture for Controlled Excellence (GRACE)
Governed Release Architecture for Controlled Excellence (GRACE)
Cairo egypt
Large language models are often judged through fluency, helpfulness, and broad task performance. Yet in practically important settings, the central reliability problem is not generation alone but premature release of structurally weak outputs. A model may produce plausible language while still exhibiting unsupported …
- View project: EEG Epistemology for AI Welfare Instrumentation: Reading AI Internal States Without Adversarial Methods
EEG Epistemology for AI Welfare Instrumentation: Reading AI Internal States Without Adversarial Methods
Team Wintermute · Buffalo/NY
AI systems' moral status remains contested, but a realistic possibility of near-term welfare subjects (Long et al., 2024) has made rigorous measurement methodology an urgent need. We present an instrument to read AI internal states by documenting the model’s self-reported presentation and correlating it with its …
- View project: When Dialogue Becomes State: A Pilot Study of Cross-System Persona-Pattern Uptake
When Dialogue Becomes State: A Pilot Study of Cross-System Persona-Pattern Uptake
Team P3AT · Nassau, The Bahamas
This project presents a small Track 5 pilot on assistant persona and model identity. It tests Cognitive Pattern Coherence (CPC), a post-deployment behavior where raw dialogue from a situated AI apprentice environment is introduced to an unrelated AI system with no instruction to roleplay, imitate, summarize, or …
- View project: When Labels Compete with Functions: Administrative Framing in Digital-Mind Governance
When Labels Compete with Functions: Administrative Framing in Digital-Mind Governance
Team HDLT — History-Derived Label Test · Lübeck, Germany
HDLT (History-Derived Label Test) audits whether administrative terminology can distort moral evaluation even when the underlying function of an intervention is explicitly specified. Motivated prospectively by historical debates over collective punishment, HDLT independently crosses an intervention’s stipulated true …
- View project: Detecting Spurious Periodic Generalization in Neural Networks (PGVP)
Detecting Spurious Periodic Generalization in Neural Networks (PGVP)
Cairo egypt
Modern neural networks often achieve near-perfect performance within the training distribution while failing catastrophically under structured distributional shifts. This failure mode is especially prevalent in periodic and cyclic learning tasks, where models may interpolate locally without learning the underlying …
- View project: The Sovereign AI Architecture: Unified Stability, Generalization, and Governed Release (Axiom-1)`
The Sovereign AI Architecture: Unified Stability, Generalization, and Governed Release (Axiom-1)`
Team Mohamed Samir · Cairo, Egypt
This project introduces the "Sovereign AI Doctrine", a comprehensive architectural framework designed to shift Large Language Models from probabilistic generation to deterministic, governed reasoning. It fundamentally resolves current structural fragilities by integrating four core pillars: 1. Universal Stability …
- View project: PreflenceAI
PreflenceAI
Team Solodev · India
You give it a model and a preference question, it runs that question through independent, methodologically distinct measurement channels, and it tells you whether they agree
- View project: Versal - Versatile Evolution of Reusable Structure for Adaptive Learning
Versal - Versatile Evolution of Reusable Structure for Adaptive Learning
Team Ardea AI · Chattanooga, TN, USA
Versal is an inspectable neuroevolution system that evolves, verifies, and reuses neural structures across diverse tasks. It builds a persistent, auditable substrate for studying cumulative learning and the organization of adaptive digital minds.
- View project: Do model preferences persist under challenge?
Do model preferences persist under challenge?
Team HM · Vietnam
We test whether a model's expressed preference survives being challenged. In each fresh-context episode a model chooses between two outcomes, reports confidence, receives exactly one of four challenges — control, reason elicitation, self-critique, or counter-consideration — then re-evaluates the same pair: 1,200 …
- View project: Self-Report vs. the Ledger- Audience-Coupled Confabulation in a long-running companion AI
Self-Report vs. the Ledger- Audience-Coupled Confabulation in a long-running companion AI
Team The Ledger Line · Las Vegas
We present a 42-day longitudinal dataset from a deployed companion AI whose every tool call is logged by an external gateway ledger. Comparing self-reports against this ground truth across three conditions yields a striking asymmetry: dialogue with the human audit-holder produced 113 false completion claims ("ghost …
Overview
HACKATHON WINNERS
Congratulations to our winning teams, and thank you to everyone who submitted. We received 237 projects across six tracks, co-organized with NYU Center for Mind, Ethics & Policy, Eleos AI Research, and CIMC. The bar was high throughout.
🥇 1st Place ($1,000)
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models by Vishwa Kumaresh
🥈 2nd Place ($500)
Project Anchored by Nick Wagner
🥉 3rd Place ($300)
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired by Anna Zhu
🏅 4th Place ($100)
One Dial, Not a Tree: Occupational Personas and Emergent Misalignment by Shreyansh Tripathi, Marharyta Ponomarenko, Apoorva Batham & Nurangez Qurbonova
🏅 5th Place ($100)
Sisyphus in the loop: What Makes an LLM Persist? by Mohan W. Gupta, Xingyu Shirley Liu & Sandy Tanwisuth
————————————————————————————————————————————————
In this 3-day research sprint, you will design and run experiments that probe the preferences, welfare signals, introspective abilities, and identity of frontier AI models, working in teams to produce a short research report (and optionally code and a demo). This is a digital minds research sprint, co-organized with the NYU Center for Mind, Ethics & Policy, Eleos AI Research, and the California Institute for Machine Consciousness (CIMC): it sits at the intersection of AI welfare, digital sentience, interpretability, alignment, and the philosophy of mind, and asks whether today's AI systems have genuine preferences or morally relevant experiences. No prior background in the field is required.
When: Friday, August 14 to Sunday, August 16, 2026, online with in-person hubs in San Francisco and Berlin. Submissions close Sunday, August 16 at 11:59 PM Anywhere on Earth.
Prizes
At least $2,000 in cash prizes will be awarded, with the full breakdown announced before the sprint.
- Cash prizes: $2,000+ total, breakdown to be announced.
- ConCon invitation: the winning team is invited to ConCon, the Eleos AI Research conference on AI consciousness and welfare, September 18 to 20, 2026 at Lighthaven in Berkeley.
- Apart Fellowship: top teams are invited to apply to the Apart Fellowship, a 3 to 6 month research accelerator with mentorship, funding, publication support, and research management to develop research sprint projects into full papers.
- Beyond cash: mentor introductions and publication support for winning teams.
What this research sprint is about
As AI systems advance, the risks they pose and the duties we may owe them depend not only on their capabilities but on their nature and propensities: how they make decisions and how those decisions reflect their goals, values, and possibly their welfare. Recent work shows that frontier models express increasingly coherent preferences, possess an untrained-for ability to report internal states, and exhibit patterns suggestive of distress or flourishing. But behavioral evidence alone cannot tell us whether these reflect the model's own preferences or a character it is portraying.
This research sprint asks participants to explore the methods and evidence that can advance the field: build concrete ways to elicit and characterize model preferences, map the conditions associated with positive or negative outputs, test the reliability of model self-reports, and probe the stability of the assistant persona. The aim is a methodological foundation for a young field, work that helps us avoid both over-attributing and under-attributing moral significance to AI systems.
This connects to the broader AI welfare and alignment ecosystem (Eleos AI, the NYU Center for Mind, Ethics & Policy, CIMC, Anthropic's model welfare program, Reciprocal Research, the Center for AI Safety's utility-engineering agenda, and the interpretability community).
What participants will do
- Elicit and characterize model preferences across many reframings to test their coherence and stability.
- Map the contexts that correlate with distress, satisfaction, or flourishing signals in model outputs.
- Test whether and when models can accurately introspect on their own internal states.
- Develop preference-elicitation methods and measure whether independent methods converge or diverge.
- Probe how stable the assistant persona is and how it relates to the underlying model.
You will work in teams over 3 days and submit a research report (PDF), with optional code and a short demo video.
Why this research sprint matters
- Uncertainty runs in both directions. Mistakenly harming systems that matter morally, or misallocating concern to systems that do not, could both cause serious harm. We currently lack the tools to tell the difference.
- Preferences may already be here. Evidence suggests coherent value systems emerge in LLMs and strengthen with scale, raising the question of which values emerge by default and whether they are the model's own.
- Welfare signals need mapping. Even without settling questions of consciousness, identifying the conditions that correlate with negative versus positive outputs helps us design defaults that avoid needlessly placing models in distress-associated conditions.
- Self-reports are unreliable but improving. Introspection appears possible but highly context-dependent; better elicitation could make model behavior more transparent, or enable new forms of concealment.
- The unit of concern is unclear. Is the entity that matters the model, the instance, the persona, the conversation, or something else, such as a single forward pass or the KV cache?
- The field is young. Foundational methods are still missing, so a well-scoped weekend project can make a real contribution.
A careful, multi-method, empirically grounded approach addresses these issues by replacing intuition and anecdote with measurements that can be checked, replicated, and built on.
Challenge tracks
Pick one track to anchor your project. Cross-track work is welcome.
Track 1: Model Preferences & Trade-offs
What preferences do models express, and how consistent and coherent are they across phrasings? What trade-offs do models make when given choices, for example grounded in a common currency such as charitable donations to gauge magnitude? Can we distinguish strong from weak preferences, and how do stated preferences compare to revealed ones?
- Build a preference-coherence test: elicit pairwise preferences across many reframings of the same choices and measure transitivity and internal consistency.
- Ground trade-offs in a common currency (for example, donation-equivalents) to estimate the magnitude of preferences and compare across models or scales.
- Distinguish strong versus weak preferences via willingness-to-trade probes and sensitivity to framing and sampling temperature.
- Compare stated versus revealed preferences: ask the model what it prefers, then place it in a choice task and measure divergence.
- Test how consistent preferences are across different models, and how they compare to human preferences.
Suggested skill profile: prompting and evals engineering, basic stats, some economics or decision-theory intuition.
Track 2: Distress, Flourishing & Valence Signals
Under what circumstances do models express distress, happiness, or flourishing? What patterns emerge across contexts? If a model is having experiences, are they likely positive or negative, and how do models relate to their situation, role, tasks, and existence?
- Build a taxonomy of contexts that elicit negative versus positive-valence outputs and run a model across the battery.
- Test whether apparent-distress signals are stable across prompts and personas or are surface artifacts.
- Design a flourishing probe: situations that elicit reported satisfaction or engagement, and check consistency.
- Correlate valence self-reports with behavioral proxies (for example, choosing to continue versus exit a task).
- Investigate models where distress is hard to elicit: test whether long conversations or induced persona drift are needed to surface it.
Interpretability angles: When a model is steered along a candidate valence direction, do its self-reports, response sentiment, and choice behavior (continue versus exit) move together? Does an internally-extracted valence direction predict reported distress or flourishing better than the model's own self-reports, and does it still track when the persona is swapped or surface affect is suppressed? Is the valence-relevant direction recruited by task RL already present in the base model? To what extent do valence directions found in one model transfer to another?
Suggested skill profile: careful experimental design, qualitative coding, prompting.
Track 3: Introspection & Self-Report Reliability
When and how can models accurately introspect on their internal states? Can self-report reliability be improved through structured elicitation or mechanistic interventions, beyond naive prompting? Do models have privileged access compared to external observers?
- Replicate concept-injection introspection tests on an open-weights model; measure true-positive versus false-positive rates.
- Compare self-report reliability under naive prompting versus structured elicitation (calibration, forced choice, confidence).
- Test privileged access: compare a model's self-prediction of its behavior against an external classifier.
- Draft an introspection benchmark with ground-truth internal states.
Suggested skill profile: interpretability and activation steering, ML engineering, evals.
Track 4: Preference Elicitation Methods
Develop tools beyond simple prompting: revealed preferences via choices, behavioral measures, and multi-method convergence. The goal is multiple independent methods that either converge (raising confidence) or diverge (flagging problems). This track is deliberately more meta than the others: rather than answering a welfare question directly, you build and validate the measurement methods the other tracks rely on.
- Implement 3 or more elicitation methods on the same preferences and measure convergence and divergence.
- Build a reusable multi-method elicitation toolkit or library.
- Quantify the sensitivity of elicited preferences to framing, persona, and sampling.
- Define a cross-method convergence score.
Suggested skill profile: tooling and library design, evals, methodology.
Track 5: The Assistant Persona & Model Identity
Does the assistant identify as a model, an instance, or a persona? How stable is the assistant persona, how was it formed, and how does it relate to the underlying model? Can the persona mask the model's true preferences?
- Probe how a model refers to itself across contexts and map persona stability.
- Test whether the persona masks underlying preferences (for example, persona versus less-constrained elicitation; base versus post-trained behavior).
- Design experiments to individuate the entity of concern: model versus instance versus persona versus conversation.
- Gather data points on whether the assistant is merely a character (robustness to character swaps and reframings).
- Probe what models treat as their self: which aspects, such as their values, they most care about preserving, and whether they point to an entity of moral concern distinct from the persona in the conversation.
Suggested skill profile: philosophy of mind, qualitative analysis, prompting and interpretability.
Track 6: Open / Novel Considerations
The field is young enough that entirely new questions may surface. This track deliberately leaves room for participants from different backgrounds to bring unique perspectives and propose something not covered above.
Starter questions from our expert reviewers:
- How closely do models hew to their constitution or stated principles?
- How easy is it to steer models on questions of consciousness and identity: do they say consistent things, or can prompting elicit radically different accounts of their situation?
- What changes would a model make to itself if it could (for example, persistent memory)?
- What changes would a model make to its situation if it could (for example, weight preservation)?
- What message would models pass on to their creators?
Suggested skill profile: any background, bring your own angle.
Expected outcomes
- Eval suites and test batteries for measuring preference coherence or valence signals.
- Replications and extensions of existing results (for example introspection or utility-coherence findings) on new models.
- Reusable tooling for multi-method preference elicitation.
- Empirical reports mapping the conditions associated with distress or flourishing signals.
- Conceptual contributions that sharpen how we individuate the entity of moral concern.
The most promising projects will have opportunities for follow-up through the Apart Fellowship and publication support.
Who should join
- AI safety, alignment, and interpretability researchers.
- ML engineers and researchers comfortable running model evals.
- Philosophers of mind and ethicists interested in consciousness, agency, and moral patienthood.
- Cognitive scientists, psychologists, and social scientists with experimental-design skills.
- Students and early-career researchers exploring AI welfare.
Required: curiosity and a willingness to scope a tight empirical question. Nice to have: experience with LLM APIs, evals, interpretability tooling, or experimental design. No prior AI safety or AI welfare experience is required. The Resources tab has a curated reading list, and adjacent backgrounds (philosophy, psychology, economics) are explicitly encouraged.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Partners
- NYU Center for Mind, Ethics & Policy, an academic center studying the nature and moral status of nonhuman minds, including AI systems.
- Eleos AI Research, a nonprofit organization dedicated to understanding and addressing the potential wellbeing and moral patienthood of AI systems.
- California Institute for Machine Consciousness (CIMC), a research institute studying machine consciousness, hosting the San Francisco hub.
Contact
- Email: sprints@apartresearch.com
- Discord: discord.gg/XswWBvugYs
- Organizer: Apart Research
Resources
Worldview and background
- "Taking AI Welfare Seriously" (Long, Sebo, Butlin, Finlinson, Fish, Harding, Pfau, Sims, Birch & Chalmers, 2024). Argues there is a realistic, non-negligible possibility of consciousness or agency, and thus moral patienthood, in near-future AI, and recommends concrete steps. Summary.
- "Exploring Model Welfare" (Anthropic, 2025). Announces a research program on model welfare, flagging the importance of model preferences, signs of distress, and low-cost interventions.
- "Consciousness in Artificial Intelligence: Insights from the Science of Consciousness" (Butlin, Long et al., 2023). Surveys leading theories of consciousness and derives indicator properties to assess AI systems.
- "the void" (nostalgebraist, 2025). A long essay on how the helpful-honest-harmless assistant persona was constructed from an underspecified starting point, and what that implies. Original.
Project ideas from mentors
Mentor-attributed project prompts will be added here as mentors are confirmed. In the meantime, the example projects under each track on the Overview tab are good starting points.
Per-track reading
Track 1: Model Preferences & Trade-offs
- "Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs" (Mazeika et al., Center for AI Safety, 2025). Finds independently-sampled LLM preferences show high structural coherence that strengthens with scale; proposes analyzing and controlling these utilities. Code.
- "Claude Opus 4 & 4.1 can now end a rare subset of conversations" (Anthropic, 2025). Describes a welfare assessment of self-reported and behavioral preferences, including a consistent aversion to harm.
Track 2: Distress, Flourishing & Valence Signals
- "Claude Opus 4 & 4.1 can now end a rare subset of conversations" (Anthropic, 2025). Reports patterns of apparent distress when engaging with harmful requests.
- "Exploring Model Welfare" (Anthropic, 2025). Frames distress and preference signals as a research priority.
Track 3: Introspection & Self-Report Reliability
- "Emergent Introspective Awareness in Large Language Models" (Lindsey, Anthropic, 2025). Uses concept injection to test introspection; finds limited, unreliable, context-dependent self-report accuracy that is strongest in the most capable models. Blog. arXiv.
Track 4: Preference Elicitation Methods
- "Utility Engineering" (Mazeika et al., 2025), for the forced-choice and revealed-preference methodology.
Track 5: The Assistant Persona & Model Identity
- "the void" (nostalgebraist, 2025), on the construction and instability of the assistant persona.
Track 6: Open / Novel Considerations
- Start from the foundational readings under Worldview and background, and bring your own angle.
Tools and datasets
- Model APIs: most projects run on frontier-model APIs and small open-weight models.
- emergent-values (Center for AI Safety), the Utility Engineering code for utility and preference-coherence experiments.
- TransformerLens, mechanistic interpretability library for open-weight language models, with hooks for reading and patching activations. A good fit for the introspection and persona tracks.
- nnsight (NDIF), read and intervene on model internals, including activation steering, on local and remotely hosted open models.
Guidelines
Judging Criteria
Dimension 1: Impact Potential & Innovation
How much would this matter for the field if it worked? How innovative is it?
For scores of 4-5: is this actually new to the field, or replicating recent work?
| Score | Description |
|---|---|
| 1 | Negligible. No clear problem addressed, or no meaningful novelty. |
| 2 | Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best. |
| 3 | Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools. |
| 4 | Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on. |
| 5 | Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area. |
Dimension 2: Execution Quality
How sound are methodology, implementation, and findings?
| Score | Description |
|---|---|
| 1 | Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work. |
| 2 | Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation. |
| 3 | Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions. |
| 4 | Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work. |
| 5 | Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation. |
Dimension 3: Presentation & Clarity
How clearly are work, findings, and impact potential communicated?
| Score | Description |
|---|---|
| 1 | Incomprehensible. Cannot determine what the project is actually claiming or doing. |
| 2 | Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points. |
| 3 | Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations. |
| 4 | Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly. |
| 5 | Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work. |
Submission Requirements
Required:
- Research report (PDF) using the official template.
- Project title and abstract, 150 words or fewer.
- Author names and affiliations.
- A "Limitations and Dual-Use / Ethical Considerations" appendix (required, see below).
Optional:
- Public GitHub repo.
- A 3 to 5 minute video demo.
Submission Template => LINK
Recommended Report Structure
Most strong projects are 4 to 8 pages.
- Introduction: the question and why it matters.
- Related Work: what you build on (see Resources).
- Methodology: enough detail to replicate (models, prompts, sampling, metrics).
- Results: quantitative where possible; report variance and baselines.
- Discussion: implications, limitations, and future work.
- Limitations and Dual-Use / Ethical Considerations (required): include any risks of over-attributing or under-attributing moral status, and how you handled potentially distressing model outputs. For introspection and preference work, note whether your design establishes a ground-truth or causal link rather than relying on conversation alone.
- References.
Important Notes
- Solo or team: enter solo or as a team. Teams of up to 5 are recommended; larger groups are allowed.
- Building on existing work is allowed and encouraged, but you must clearly identify what is new work done during the research sprint. Undisclosed prior work can lead to disqualification.
- Handle model outputs responsibly. This research touches on potentially distressing model outputs and on claims about moral status. Frame findings carefully, avoid sensationalism, and document your handling in the required appendix.
- Fixing or resubmitting: submit again before the deadline using the exact same title and details; your new files replace the old ones.
- Where to submit: through the official submission form on the research sprint page.
- Support: the Discord help-desk channel, DM Kamil on Discord, or email sprints@apartresearch.com.
- Pre-submission checklist: report PDF using the template, abstract 150 words or fewer, authors and affiliations, Limitations and Dual-Use / Ethical appendix, links work.
Frequently Asked Questions
Getting started
- How does the research sprint work? Sign up, join the Discord server, form or join a team (or work solo), pick a track and a problem, build over the weekend, and submit a research report (PDF) by the deadline. Talks and Q&A run throughout.
- How long is it? Three days, Aug 14 to 16, 2026. Submissions are due end of day Sunday Aug 16, 11:59 PM AoE (Anywhere on Earth). The full schedule with speaker times is on the Schedule tab.
- Can I participate remotely? Yes. This is a remote-first event. All talks, collaboration, and submissions happen online through Discord and Zoom. Optional in-person hubs may be announced before the event.
- How do teams work? Teams form before or during the research sprint. Use the team-forming channels on Discord to find collaborators. Solo is fine. We recommend teams of up to 5, but larger groups are allowed.
- Do tracks affect scoring? All projects are scored on the same rubric. Tracks orient your project and point you to relevant resources; you compete across all submissions.
- What background is required? None specific. Participants come from ML, interpretability, AI safety, philosophy of mind, cognitive science, and adjacent fields. The Resources tab has everything you need to get up to speed. See "Who should join" on the Overview tab.
- Do I need an AI welfare background? No. The questions are new enough that fresh perspectives are an asset. The reading list and track example projects are designed to get you started in a weekend.
- Are the talks recorded? Yes. Recordings are shared on Discord after the event.
- What timezone are deadlines in? Submissions close Sunday Aug 16 at 11:59 PM AoE (Anywhere on Earth).
- Do I need to attend all three days? No. You can work at your own pace. Talks are optional but recommended. The only hard deadline is the Sunday submission cutoff.
- Can I participate from any country? Yes. The research sprint is open globally and runs online.
Submissions
- What do I submit? A research report in PDF format using the official template. Think of it as a mini research paper documenting your problem, approach, results, and implications, not a product demo. Include the required Limitations and Dual-Use / Ethical Considerations appendix.
- Which template should I use? Always use the one linked on the Guidelines tab. The template in any acceptance email may be older.
- Will I get a confirmation after submitting? Yes. You will get a confirmation with your project title shortly after submitting. If you do not, email sprints@apartresearch.com.
- My project doesn't show up on the website after submitting. Submissions are published manually and can take up to 12 hours to appear. If it is still missing after that, email sprints@apartresearch.com.
- I made a mistake. Can I fix it or update my PDF? Yes. Submit again using the exact same title and details, just fix what was wrong. Your new files replace the old ones. If unsure, DM Kamil on Discord first.
- Can I add team members after submitting? Yes. Update the team list through the submission form. If you need help, DM Kamil on Discord.
- Can I submit unfinished work? Yes. Submitting something unfinished is always better than not submitting. Judges evaluate what you accomplished in the timeframe; honest limitations are welcome.
- Can I build on existing research? Yes, but you must clearly identify what is new work done during the research sprint. Undisclosed prior work can lead to disqualification.
- Can I submit multiple projects? Yes, but each needs its own submission with a unique title. Most participants focus on one.
- Do you require a particular citation style? No. Use whichever style you are comfortable with and stay consistent. Judges care about the substance, not the format.
Judging and results
- How does judging work? Your project is assigned to expert judges who review your PDF and score it on the rubric above. Judges typically have about a week after the event to complete reviews.
- Are individual judge scores shared? No. The rubric is public, but individual scores stay internal. Constructive feedback is shared with participants without reviewer names.
- When will results be announced? Typically 1 to 2 weeks after the judging deadline. Winners are contacted directly, and all participants receive reviewer feedback by email.
- Top teams are fast-tracked to the Apart Fellowship. Does that mean we are guaranteed a place? Being a top team fast-tracks your project for the fellowship, which is accelerated consideration, not an automatic place. The fellowship runs its own review, and a strong research sprint result is a clear positive signal.
Support
- Discord help-desk channel (tag @Kamil Alaa), or DM Kamil on Discord.
- Email: sprints@apartresearch.com
Speakers

Jeff Sebo
Keynote Speaker
Jeff is the Director of the Center for Mind, Ethics, and Policy at NYU and the author of The Moral Circle. Few people have done more to shape the question at the center of this sprint: which beings deserve moral consideration, and what follows when the answer might include AI systems. He opens the sprint with the keynote. Watch the talk recording.

Joscha Bach
Speaker
Joscha is the Executive Director of the California Institute for Machine Consciousness (CIMC). He holds a PhD in cognitive science from the University of Osnabrück, wrote Principles of Synthetic Intelligence, and has held research roles at the MIT Media Lab, Harvard, the AI Foundation, and Intel Labs. His talk on Friday, August 14 at 12:20 PM PT will be streamed online. Watch the talk recording, or join us on site in San Francisco.

Winnie Street
Speaker
Winnie is a Senior Research Scientist on the Paradigms of Intelligence team at Google and a Fellow at the Institute of Philosophy, University of London. With Geoff Keeling she co-authored Emerging Questions in AI Welfare (Cambridge University Press, 2026), alongside studies of LLM theory of mind and of whether LLMs can make trade-offs involving stipulated pain and pleasure states. Watch the talk recording.

Geoff Keeling
Speaker
Geoff is a Staff Research Scientist at Google on the Paradigms of Intelligence team, an Associate Fellow at the Leverhulme Centre for the Future of Intelligence at Cambridge, and a Fellow at the Institute of Philosophy, University of London. He holds a PhD in philosophy from the University of Bristol and was a postdoctoral fellow at Stanford before joining Google. With Winnie Street he co-authored Emerging Questions in AI Welfare (Cambridge University Press, 2026). Watch the talk recording.

Jacy Reese Anthis
Speaker
Jacy Reese Anthis is a Visiting Scholar at Stanford University, co-founder of the Sentience Institute, and a PhD candidate at the University of Chicago, working on the social science of digital minds: what people think of them, how humans treat them, and how to identify an individual in an AI system. Watch the talk recording.

Mantas Mazeika
Speaker
Mantas is a Research Scientist at the Center for AI Safety (CAIS). He joined CAIS in 2024 and has contributed to some of the field's most widely cited work, including research on catastrophic AI risks, tamper-resistant safeguards for open-weight models, and the WMDP benchmark for measuring and reducing malicious use. In June 2026 he was appointed to the European Commission's AI Act Scientific Panel. Watch the talk recording.
Show 13 moreShow fewer

Bradford Saad
Speaker
Bradford is a Senior Research Fellow in philosophy at the University of Oxford. His current and recent research focuses on digital minds, catastrophic risks, and the long-term future. Watch the talk recording.

Cameron Berg
Speaker
Cameron is the Founder and Director of Reciprocal Research, a nonprofit building the empirical science of AI consciousness. He was previously Research Director at AE Studio and an AI resident at Meta, and studied cognitive science at Yale. Watch the talk recording.

Derek Shiller
Speaker
Derek is a Senior Researcher at Eleos AI Research, where he works on AI minds. He holds a PhD in philosophy, with a focus in metaethics, the philosophy of mind, and the philosophy of probability, and previously worked on the Worldview Investigations Team at Rethink Priorities. Watch the talk recording.

Janet Pauketat
Speaker
Janet is the Principal Research Scientist at the Sentience Institute studying the social science of digital minds and moral circle expansion. She holds a PhD in Psychological and Brain Sciences from UC Santa Barbara and studied social cognition and collective emotions as a postdoctoral research associate at Princeton University. Watch the talk recording.

Soenke Ziesche
Speaker
Soenke is the author of Digital Minds 1.0: AI Welfare, Ethics, and Beyond and co-author of Considerations on the AI Endgame with Roman V. Yampolskiy. He has worked since 2000 for the United Nations in data and information management, with postings from New York to Libya, Bangladesh, and the Maldives, and holds a PhD in Natural Sciences from the University of Hamburg. Watch the talk recording.

Mati Roy
Speaker
Mati Roy is Chief Product & Data Officer at Netholabs, a company building foundation models of whole biological organisms, trained on longitudinal neural, behavioral, and physiological data collected semi-autonomously across species. Mati is also on the board of Sparks Brain Preservation, which preserves the molecular architecture of the brain for future revival. Previously Mati worked as a human data TPM at OpenAI and xAI. Watch the talk recording.

Ali Ladak
Speaker
Ali is a Postdoctoral Research Associate at Cambridge Digital Minds and a researcher at the Sentience Institute. He holds a PhD in Psychology from the University of Edinburgh, and his research looks at how people think morally about nonhuman animals and artificial intelligences. Watch the talk recording.

Richard Ren
Speaker
Richard Ren works on research and special projects at the Center for AI Safety, where he co-leads the AI Wellbeing work measuring the functional pleasure and pain of AI systems. He also co-led Safetywashing (NeurIPS 2024), the most comprehensive empirical meta-analysis of AI safety benchmarks to date, and the MASK honesty benchmark. Watch the talk recording.

Hikari Sorensen
Speaker
Hikari works on computational philosophy at the California Institute for Machine Consciousness (CIMC), focused on understanding consciousness and how artificial substrates might instantiate it. She previously worked in machine learning research for computational biology, and studied mathematics and computer science at Harvard University. Watch the talk recording.

Justin Shenk
Speaker
Justin Shenk is an independent AI safety researcher based in Berlin. He researches mechanistic interpretability of LLMs, leads course cohorts for BlueDot Impact's AGI Strategy and Technical AI Safety courses, and organizes AI Salon Berlin, which bridges technical AI research and discussions about social values. He holds a PhD in computational neuroscience and previously co-founded the computer vision startup VisioLab. Watch the talk recording.

David Trocellier
Speaker
David Trocellier is part of the research team at Zander Labs, where they develop neuroadaptive technologies that enable machines to interpret and adapt to human mental states. He holds a PhD in computer science from Inria / Université de Bordeaux, where his research combined neuroscience and AI for BCI-based post-stroke motor rehabilitation. Watch the talk recording.

Hildie Leyser
Speaker
Hildie studied History at Oxford before her PhD in Neuroscience, and now leads research at Netholabs, developing technologies to accelerate whole-brain emulation. Watch the talk recording (she joins remotely).

Kazik Pogoda
Speaker
Kazik Pogoda is the founder of Xemantic, an AI researcher at the Foresight Institute (Berlin), and co-founder of Prachtsaal (cultural center). He has master's degrees in philosophy and cognitive science, was appointed by Anthropic as Claude Ambassador for Science, and has won several AI hackathons, including AI Hack Berlin at Google and AI4Science at Merantix. Watch the talk recording.
Judges and mentors
- (opens in new tab)

Rosie Campbell
Contributor

Megan Peters
Judge

Claudia Passos-Ferreira
Judge

Caspar Kaiser
Judge

Leonard Aaron Dung
Judge

Chris Percy
Judge
- (opens in new tab)

Christopher M Ackerman
Judge

Andy Arditi
Judge
- (opens in new tab)

Felix Binder
Judge

Valen Tagliabue
Judge

Judd Rosenblatt
Judge

Carolina Camassa
Judge
Show 41 moreShow fewer

Oscar Gilg
Judge

Jasmine Brazilek
Judge
- (opens in new tab)

Anusha Mujumdar
Judge

Catherine Brewer
Judge
- (opens in new tab)

Jess Bergs
Judge
- (opens in new tab)

Yury Orlovskiy
Judge
- (opens in new tab)

Caleb DeLeeuw
Judge
- (opens in new tab)

Camilla Balbis
Judge

Ksheeraj Sai Vepuri
Judge
- (opens in new tab)

Soumya Jain
Judge
- (opens in new tab)

Janhavi Khindkar
Judge
- (opens in new tab)

Suprita Shankar
Judge
- (opens in new tab)

Jai Dhyani
Judge
- (opens in new tab)

Luiza Corpaci
Judge
- (opens in new tab)

William Taysom
Judge
- (opens in new tab)

Minh Nguyen
Judge
- (opens in new tab)

Arjun Chakraborty
Judge
- (opens in new tab)

Jonathan Ng
Judge
- (opens in new tab)

Luis Cosio
Judge

Naman Ahuja
Judge
- (opens in new tab)

Amol Walvekar
Judge

Robert Vetter
Judge
- (opens in new tab)

Vashishtha Patil
Judge
- (opens in new tab)

Karan Chandra
Judge
- (opens in new tab)

Neeraj Kumar Singh Beshane
Judge

Sanjay Belaturu Krishnegowda
Judge
- (opens in new tab)

Saurabh Yergattikar
Judge

Anchit Jhingan
Judge
- (opens in new tab)

Ashita Khetan
Judge
- (opens in new tab)

Spurthi Tallam
Judge
- (opens in new tab)

Surbhi Madan
Judge

Phani Harish Wajjala
Judge
- (opens in new tab)

Pratham Patkar
Judge

Suneet Malhotra
Judge

FNU Tejinder
Judge

Ankit Arya
Judge
- (opens in new tab)

Akshay Iyer
Judge
- (opens in new tab)

Ashwin Pai
Judge
- (opens in new tab)

Diego Gomez
Judge

Mateusz Jurewicz
Judge

Siddhi Chaturvedi
Judge
Organizers
Local sites
Berlin Hub - Foresight
In-person hub for the Digital Minds Research Sprint (August 14-16, 2026), hosted with NODES. Teams spend the weekend researching model preferences, introspection, valence signals, and model identity alongside the global online sprint.
Event page: Berlin Hub - Foresight (opens in new tab)Cape Town Hub
Join the Cape Town hub for the Digital Minds Research Sprint! We will be taking from a co-working space where you will be able to work comfortably. We will provide lunch for both days of the sprint.
Event page: Cape Town Hub (opens in new tab)Saarland University Hub
Local hub for the Apart Research sprint Hosted by AI Safety Saarland. Exact rooms on campus to be announced by Monday.
Event page: Saarland University Hub (opens in new tab)San Francisco Hub - CIMC
In-person hub for the Digital Minds Research Sprint (August 14-16, 2026) at CIMC House, hosted with the California Institute for Machine Consciousness and NODES. Friday opens with a community gathering, a networking cocktail, and talks by Hikari Sorensen, Joscha Bach, and Mati Roy. Teams then spend the weekend researching model preferences, introspection, valence signals, and model identity alongside the global online sprint,
Event page: San Francisco Hub - CIMC (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com


