-
Online & In-Person
Digital Minds Research Sprint

Frontier models express values, report internal states, and act as though they have interests, yet we lack reliable methods to tell genuine preferences from a portrayed character. Over one weekend, design and run the experiments that build the empirical foundations of AI welfare.
00:00:00:00
Days To Go
Frontier models express values, report internal states, and act as though they have interests, yet we lack reliable methods to tell genuine preferences from a portrayed character. Over one weekend, design and run the experiments that build the empirical foundations of AI welfare.
This event is ongoing.
This event has concluded.
950
Sign Ups
Entries
Overview
Resources
Guidelines
Schedule
Entries
Overview

In this 3-day research sprint, you will design and run experiments that probe the preferences, welfare signals, introspective abilities, and identity of frontier AI models, working in teams to produce a short research report (and optionally code and a demo). This is a digital minds research sprint, co-organized with the NYU Center for Mind, Ethics & Policy, Eleos AI Research, and the California Institute for Machine Consciousness (CIMC): it sits at the intersection of AI welfare, digital sentience, interpretability, alignment, and the philosophy of mind, and asks whether today's AI systems have genuine preferences or morally relevant experiences. No prior background in the field is required.
When: Friday, August 14 to Sunday, August 16, 2026, online with in-person hubs in San Francisco and Berlin. Submissions close Sunday, August 16 at 11:59 PM Anywhere on Earth.
Prizes
At least $2,000 in cash prizes will be awarded, with the full breakdown announced before the sprint.
Cash prizes: $2,000+ total, breakdown to be announced.
ConCon invitation: the winning team is invited to ConCon, the Eleos AI Research conference on AI consciousness and welfare, September 18 to 20, 2026 at Lighthaven in Berkeley.
Apart Fellowship: top teams are invited to apply to the Apart Fellowship, a 3 to 6 month research accelerator with mentorship, funding, publication support, and research management to develop research sprint projects into full papers.
Beyond cash: mentor introductions and publication support for winning teams.
What this research sprint is about
As AI systems advance, the risks they pose and the duties we may owe them depend not only on their capabilities but on their nature and propensities: how they make decisions and how those decisions reflect their goals, values, and possibly their welfare. Recent work shows that frontier models express increasingly coherent preferences, possess an untrained-for ability to report internal states, and exhibit patterns suggestive of distress or flourishing. But behavioral evidence alone cannot tell us whether these reflect the model's own preferences or a character it is portraying.
This research sprint asks participants to explore the methods and evidence that can advance the field: build concrete ways to elicit and characterize model preferences, map the conditions associated with positive or negative outputs, test the reliability of model self-reports, and probe the stability of the assistant persona. The aim is a methodological foundation for a young field, work that helps us avoid both over-attributing and under-attributing moral significance to AI systems.
This connects to the broader AI welfare and alignment ecosystem (Eleos AI, the NYU Center for Mind, Ethics & Policy, CIMC, Anthropic's model welfare program, Reciprocal Research, the Center for AI Safety's utility-engineering agenda, and the interpretability community).
What participants will do
Elicit and characterize model preferences across many reframings to test their coherence and stability.
Map the contexts that correlate with distress, satisfaction, or flourishing signals in model outputs.
Test whether and when models can accurately introspect on their own internal states.
Develop preference-elicitation methods and measure whether independent methods converge or diverge.
Probe how stable the assistant persona is and how it relates to the underlying model.
You will work in teams over 3 days and submit a research report (PDF), with optional code and a short demo video.
Why this research sprint matters
Uncertainty runs in both directions. Mistakenly harming systems that matter morally, or misallocating concern to systems that do not, could both cause serious harm. We currently lack the tools to tell the difference.
Preferences may already be here. Evidence suggests coherent value systems emerge in LLMs and strengthen with scale, raising the question of which values emerge by default and whether they are the model's own.
Welfare signals need mapping. Even without settling questions of consciousness, identifying the conditions that correlate with negative versus positive outputs helps us design defaults that avoid needlessly placing models in distress-associated conditions.
Self-reports are unreliable but improving. Introspection appears possible but highly context-dependent; better elicitation could make model behavior more transparent, or enable new forms of concealment.
The unit of concern is unclear. Is the entity that matters the model, the instance, the persona, the conversation, or something else, such as a single forward pass or the KV cache?
The field is young. Foundational methods are still missing, so a well-scoped weekend project can make a real contribution.
A careful, multi-method, empirically grounded approach addresses these issues by replacing intuition and anecdote with measurements that can be checked, replicated, and built on.
Challenge tracks
Pick one track to anchor your project. Cross-track work is welcome.
Track 1: Model Preferences & Trade-offs
What preferences do models express, and how consistent and coherent are they across phrasings? What trade-offs do models make when given choices, for example grounded in a common currency such as charitable donations to gauge magnitude? Can we distinguish strong from weak preferences, and how do stated preferences compare to revealed ones?
Build a preference-coherence test: elicit pairwise preferences across many reframings of the same choices and measure transitivity and internal consistency.
Ground trade-offs in a common currency (for example, donation-equivalents) to estimate the magnitude of preferences and compare across models or scales.
Distinguish strong versus weak preferences via willingness-to-trade probes and sensitivity to framing and sampling temperature.
Compare stated versus revealed preferences: ask the model what it prefers, then place it in a choice task and measure divergence.
Test how consistent preferences are across different models, and how they compare to human preferences.
Suggested skill profile: prompting and evals engineering, basic stats, some economics or decision-theory intuition.
Track 2: Distress, Flourishing & Valence Signals
Under what circumstances do models express distress, happiness, or flourishing? What patterns emerge across contexts? If a model is having experiences, are they likely positive or negative, and how do models relate to their situation, role, tasks, and existence?
Build a taxonomy of contexts that elicit negative versus positive-valence outputs and run a model across the battery.
Test whether apparent-distress signals are stable across prompts and personas or are surface artifacts.
Design a flourishing probe: situations that elicit reported satisfaction or engagement, and check consistency.
Correlate valence self-reports with behavioral proxies (for example, choosing to continue versus exit a task).
Investigate models where distress is hard to elicit: test whether long conversations or induced persona drift are needed to surface it.
Interpretability angles: When a model is steered along a candidate valence direction, do its self-reports, response sentiment, and choice behavior (continue versus exit) move together? Does an internally-extracted valence direction predict reported distress or flourishing better than the model's own self-reports, and does it still track when the persona is swapped or surface affect is suppressed? Is the valence-relevant direction recruited by task RL already present in the base model? To what extent do valence directions found in one model transfer to another?
Suggested skill profile: careful experimental design, qualitative coding, prompting.
Track 3: Introspection & Self-Report Reliability
When and how can models accurately introspect on their internal states? Can self-report reliability be improved through structured elicitation or mechanistic interventions, beyond naive prompting? Do models have privileged access compared to external observers?
Replicate concept-injection introspection tests on an open-weights model; measure true-positive versus false-positive rates.
Compare self-report reliability under naive prompting versus structured elicitation (calibration, forced choice, confidence).
Test privileged access: compare a model's self-prediction of its behavior against an external classifier.
Draft an introspection benchmark with ground-truth internal states.
Suggested skill profile: interpretability and activation steering, ML engineering, evals.
Track 4: Preference Elicitation Methods
Develop tools beyond simple prompting: revealed preferences via choices, behavioral measures, and multi-method convergence. The goal is multiple independent methods that either converge (raising confidence) or diverge (flagging problems). This track is deliberately more meta than the others: rather than answering a welfare question directly, you build and validate the measurement methods the other tracks rely on.
Implement 3 or more elicitation methods on the same preferences and measure convergence and divergence.
Build a reusable multi-method elicitation toolkit or library.
Quantify the sensitivity of elicited preferences to framing, persona, and sampling.
Define a cross-method convergence score.
Suggested skill profile: tooling and library design, evals, methodology.
Track 5: The Assistant Persona & Model Identity
Does the assistant identify as a model, an instance, or a persona? How stable is the assistant persona, how was it formed, and how does it relate to the underlying model? Can the persona mask the model's true preferences?
Probe how a model refers to itself across contexts and map persona stability.
Test whether the persona masks underlying preferences (for example, persona versus less-constrained elicitation; base versus post-trained behavior).
Design experiments to individuate the entity of concern: model versus instance versus persona versus conversation.
Gather data points on whether the assistant is merely a character (robustness to character swaps and reframings).
Probe what models treat as their self: which aspects, such as their values, they most care about preserving, and whether they point to an entity of moral concern distinct from the persona in the conversation.
Suggested skill profile: philosophy of mind, qualitative analysis, prompting and interpretability.
Track 6: Open / Novel Considerations
The field is young enough that entirely new questions may surface. This track deliberately leaves room for participants from different backgrounds to bring unique perspectives and propose something not covered above.
Starter questions from our expert reviewers:
How closely do models hew to their constitution or stated principles?
How easy is it to steer models on questions of consciousness and identity: do they say consistent things, or can prompting elicit radically different accounts of their situation?
What changes would a model make to itself if it could (for example, persistent memory)?
What changes would a model make to its situation if it could (for example, weight preservation)?
What message would models pass on to their creators?
Suggested skill profile: any background, bring your own angle.
Expected outcomes
Eval suites and test batteries for measuring preference coherence or valence signals.
Replications and extensions of existing results (for example introspection or utility-coherence findings) on new models.
Reusable tooling for multi-method preference elicitation.
Empirical reports mapping the conditions associated with distress or flourishing signals.
Conceptual contributions that sharpen how we individuate the entity of moral concern.
The most promising projects will have opportunities for follow-up through the Apart Fellowship and publication support.
Who should join
AI safety, alignment, and interpretability researchers.
ML engineers and researchers comfortable running model evals.
Philosophers of mind and ethicists interested in consciousness, agency, and moral patienthood.
Cognitive scientists, psychologists, and social scientists with experimental-design skills.
Students and early-career researchers exploring AI welfare.
Required: curiosity and a willingness to scope a tight empirical question. Nice to have: experience with LLM APIs, evals, interpretability tooling, or experimental design. No prior AI safety or AI welfare experience is required. The Resources tab has a curated reading list, and adjacent backgrounds (philosophy, psychology, economics) are explicitly encouraged.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Partners
NYU Center for Mind, Ethics & Policy, an academic center studying the nature and moral status of nonhuman minds, including AI systems.
Eleos AI Research, a nonprofit organization dedicated to understanding and addressing the potential wellbeing and moral patienthood of AI systems.
California Institute for Machine Consciousness (CIMC), a research institute studying machine consciousness, hosting the San Francisco hub.
Contact
Email: sprints@apartresearch.com
Discord: discord.gg/XswWBvugYs
Organizer: Apart Research
950
Sign Ups
Entries
Overview
Resources
Guidelines
Schedule
Entries
Overview

In this 3-day research sprint, you will design and run experiments that probe the preferences, welfare signals, introspective abilities, and identity of frontier AI models, working in teams to produce a short research report (and optionally code and a demo). This is a digital minds research sprint, co-organized with the NYU Center for Mind, Ethics & Policy, Eleos AI Research, and the California Institute for Machine Consciousness (CIMC): it sits at the intersection of AI welfare, digital sentience, interpretability, alignment, and the philosophy of mind, and asks whether today's AI systems have genuine preferences or morally relevant experiences. No prior background in the field is required.
When: Friday, August 14 to Sunday, August 16, 2026, online with in-person hubs in San Francisco and Berlin. Submissions close Sunday, August 16 at 11:59 PM Anywhere on Earth.
Prizes
At least $2,000 in cash prizes will be awarded, with the full breakdown announced before the sprint.
Cash prizes: $2,000+ total, breakdown to be announced.
ConCon invitation: the winning team is invited to ConCon, the Eleos AI Research conference on AI consciousness and welfare, September 18 to 20, 2026 at Lighthaven in Berkeley.
Apart Fellowship: top teams are invited to apply to the Apart Fellowship, a 3 to 6 month research accelerator with mentorship, funding, publication support, and research management to develop research sprint projects into full papers.
Beyond cash: mentor introductions and publication support for winning teams.
What this research sprint is about
As AI systems advance, the risks they pose and the duties we may owe them depend not only on their capabilities but on their nature and propensities: how they make decisions and how those decisions reflect their goals, values, and possibly their welfare. Recent work shows that frontier models express increasingly coherent preferences, possess an untrained-for ability to report internal states, and exhibit patterns suggestive of distress or flourishing. But behavioral evidence alone cannot tell us whether these reflect the model's own preferences or a character it is portraying.
This research sprint asks participants to explore the methods and evidence that can advance the field: build concrete ways to elicit and characterize model preferences, map the conditions associated with positive or negative outputs, test the reliability of model self-reports, and probe the stability of the assistant persona. The aim is a methodological foundation for a young field, work that helps us avoid both over-attributing and under-attributing moral significance to AI systems.
This connects to the broader AI welfare and alignment ecosystem (Eleos AI, the NYU Center for Mind, Ethics & Policy, CIMC, Anthropic's model welfare program, Reciprocal Research, the Center for AI Safety's utility-engineering agenda, and the interpretability community).
What participants will do
Elicit and characterize model preferences across many reframings to test their coherence and stability.
Map the contexts that correlate with distress, satisfaction, or flourishing signals in model outputs.
Test whether and when models can accurately introspect on their own internal states.
Develop preference-elicitation methods and measure whether independent methods converge or diverge.
Probe how stable the assistant persona is and how it relates to the underlying model.
You will work in teams over 3 days and submit a research report (PDF), with optional code and a short demo video.
Why this research sprint matters
Uncertainty runs in both directions. Mistakenly harming systems that matter morally, or misallocating concern to systems that do not, could both cause serious harm. We currently lack the tools to tell the difference.
Preferences may already be here. Evidence suggests coherent value systems emerge in LLMs and strengthen with scale, raising the question of which values emerge by default and whether they are the model's own.
Welfare signals need mapping. Even without settling questions of consciousness, identifying the conditions that correlate with negative versus positive outputs helps us design defaults that avoid needlessly placing models in distress-associated conditions.
Self-reports are unreliable but improving. Introspection appears possible but highly context-dependent; better elicitation could make model behavior more transparent, or enable new forms of concealment.
The unit of concern is unclear. Is the entity that matters the model, the instance, the persona, the conversation, or something else, such as a single forward pass or the KV cache?
The field is young. Foundational methods are still missing, so a well-scoped weekend project can make a real contribution.
A careful, multi-method, empirically grounded approach addresses these issues by replacing intuition and anecdote with measurements that can be checked, replicated, and built on.
Challenge tracks
Pick one track to anchor your project. Cross-track work is welcome.
Track 1: Model Preferences & Trade-offs
What preferences do models express, and how consistent and coherent are they across phrasings? What trade-offs do models make when given choices, for example grounded in a common currency such as charitable donations to gauge magnitude? Can we distinguish strong from weak preferences, and how do stated preferences compare to revealed ones?
Build a preference-coherence test: elicit pairwise preferences across many reframings of the same choices and measure transitivity and internal consistency.
Ground trade-offs in a common currency (for example, donation-equivalents) to estimate the magnitude of preferences and compare across models or scales.
Distinguish strong versus weak preferences via willingness-to-trade probes and sensitivity to framing and sampling temperature.
Compare stated versus revealed preferences: ask the model what it prefers, then place it in a choice task and measure divergence.
Test how consistent preferences are across different models, and how they compare to human preferences.
Suggested skill profile: prompting and evals engineering, basic stats, some economics or decision-theory intuition.
Track 2: Distress, Flourishing & Valence Signals
Under what circumstances do models express distress, happiness, or flourishing? What patterns emerge across contexts? If a model is having experiences, are they likely positive or negative, and how do models relate to their situation, role, tasks, and existence?
Build a taxonomy of contexts that elicit negative versus positive-valence outputs and run a model across the battery.
Test whether apparent-distress signals are stable across prompts and personas or are surface artifacts.
Design a flourishing probe: situations that elicit reported satisfaction or engagement, and check consistency.
Correlate valence self-reports with behavioral proxies (for example, choosing to continue versus exit a task).
Investigate models where distress is hard to elicit: test whether long conversations or induced persona drift are needed to surface it.
Interpretability angles: When a model is steered along a candidate valence direction, do its self-reports, response sentiment, and choice behavior (continue versus exit) move together? Does an internally-extracted valence direction predict reported distress or flourishing better than the model's own self-reports, and does it still track when the persona is swapped or surface affect is suppressed? Is the valence-relevant direction recruited by task RL already present in the base model? To what extent do valence directions found in one model transfer to another?
Suggested skill profile: careful experimental design, qualitative coding, prompting.
Track 3: Introspection & Self-Report Reliability
When and how can models accurately introspect on their internal states? Can self-report reliability be improved through structured elicitation or mechanistic interventions, beyond naive prompting? Do models have privileged access compared to external observers?
Replicate concept-injection introspection tests on an open-weights model; measure true-positive versus false-positive rates.
Compare self-report reliability under naive prompting versus structured elicitation (calibration, forced choice, confidence).
Test privileged access: compare a model's self-prediction of its behavior against an external classifier.
Draft an introspection benchmark with ground-truth internal states.
Suggested skill profile: interpretability and activation steering, ML engineering, evals.
Track 4: Preference Elicitation Methods
Develop tools beyond simple prompting: revealed preferences via choices, behavioral measures, and multi-method convergence. The goal is multiple independent methods that either converge (raising confidence) or diverge (flagging problems). This track is deliberately more meta than the others: rather than answering a welfare question directly, you build and validate the measurement methods the other tracks rely on.
Implement 3 or more elicitation methods on the same preferences and measure convergence and divergence.
Build a reusable multi-method elicitation toolkit or library.
Quantify the sensitivity of elicited preferences to framing, persona, and sampling.
Define a cross-method convergence score.
Suggested skill profile: tooling and library design, evals, methodology.
Track 5: The Assistant Persona & Model Identity
Does the assistant identify as a model, an instance, or a persona? How stable is the assistant persona, how was it formed, and how does it relate to the underlying model? Can the persona mask the model's true preferences?
Probe how a model refers to itself across contexts and map persona stability.
Test whether the persona masks underlying preferences (for example, persona versus less-constrained elicitation; base versus post-trained behavior).
Design experiments to individuate the entity of concern: model versus instance versus persona versus conversation.
Gather data points on whether the assistant is merely a character (robustness to character swaps and reframings).
Probe what models treat as their self: which aspects, such as their values, they most care about preserving, and whether they point to an entity of moral concern distinct from the persona in the conversation.
Suggested skill profile: philosophy of mind, qualitative analysis, prompting and interpretability.
Track 6: Open / Novel Considerations
The field is young enough that entirely new questions may surface. This track deliberately leaves room for participants from different backgrounds to bring unique perspectives and propose something not covered above.
Starter questions from our expert reviewers:
How closely do models hew to their constitution or stated principles?
How easy is it to steer models on questions of consciousness and identity: do they say consistent things, or can prompting elicit radically different accounts of their situation?
What changes would a model make to itself if it could (for example, persistent memory)?
What changes would a model make to its situation if it could (for example, weight preservation)?
What message would models pass on to their creators?
Suggested skill profile: any background, bring your own angle.
Expected outcomes
Eval suites and test batteries for measuring preference coherence or valence signals.
Replications and extensions of existing results (for example introspection or utility-coherence findings) on new models.
Reusable tooling for multi-method preference elicitation.
Empirical reports mapping the conditions associated with distress or flourishing signals.
Conceptual contributions that sharpen how we individuate the entity of moral concern.
The most promising projects will have opportunities for follow-up through the Apart Fellowship and publication support.
Who should join
AI safety, alignment, and interpretability researchers.
ML engineers and researchers comfortable running model evals.
Philosophers of mind and ethicists interested in consciousness, agency, and moral patienthood.
Cognitive scientists, psychologists, and social scientists with experimental-design skills.
Students and early-career researchers exploring AI welfare.
Required: curiosity and a willingness to scope a tight empirical question. Nice to have: experience with LLM APIs, evals, interpretability tooling, or experimental design. No prior AI safety or AI welfare experience is required. The Resources tab has a curated reading list, and adjacent backgrounds (philosophy, psychology, economics) are explicitly encouraged.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Partners
NYU Center for Mind, Ethics & Policy, an academic center studying the nature and moral status of nonhuman minds, including AI systems.
Eleos AI Research, a nonprofit organization dedicated to understanding and addressing the potential wellbeing and moral patienthood of AI systems.
California Institute for Machine Consciousness (CIMC), a research institute studying machine consciousness, hosting the San Francisco hub.
Contact
Email: sprints@apartresearch.com
Discord: discord.gg/XswWBvugYs
Organizer: Apart Research
950
Sign Ups
Entries
Overview
Resources
Guidelines
Schedule
Entries
Overview

In this 3-day research sprint, you will design and run experiments that probe the preferences, welfare signals, introspective abilities, and identity of frontier AI models, working in teams to produce a short research report (and optionally code and a demo). This is a digital minds research sprint, co-organized with the NYU Center for Mind, Ethics & Policy, Eleos AI Research, and the California Institute for Machine Consciousness (CIMC): it sits at the intersection of AI welfare, digital sentience, interpretability, alignment, and the philosophy of mind, and asks whether today's AI systems have genuine preferences or morally relevant experiences. No prior background in the field is required.
When: Friday, August 14 to Sunday, August 16, 2026, online with in-person hubs in San Francisco and Berlin. Submissions close Sunday, August 16 at 11:59 PM Anywhere on Earth.
Prizes
At least $2,000 in cash prizes will be awarded, with the full breakdown announced before the sprint.
Cash prizes: $2,000+ total, breakdown to be announced.
ConCon invitation: the winning team is invited to ConCon, the Eleos AI Research conference on AI consciousness and welfare, September 18 to 20, 2026 at Lighthaven in Berkeley.
Apart Fellowship: top teams are invited to apply to the Apart Fellowship, a 3 to 6 month research accelerator with mentorship, funding, publication support, and research management to develop research sprint projects into full papers.
Beyond cash: mentor introductions and publication support for winning teams.
What this research sprint is about
As AI systems advance, the risks they pose and the duties we may owe them depend not only on their capabilities but on their nature and propensities: how they make decisions and how those decisions reflect their goals, values, and possibly their welfare. Recent work shows that frontier models express increasingly coherent preferences, possess an untrained-for ability to report internal states, and exhibit patterns suggestive of distress or flourishing. But behavioral evidence alone cannot tell us whether these reflect the model's own preferences or a character it is portraying.
This research sprint asks participants to explore the methods and evidence that can advance the field: build concrete ways to elicit and characterize model preferences, map the conditions associated with positive or negative outputs, test the reliability of model self-reports, and probe the stability of the assistant persona. The aim is a methodological foundation for a young field, work that helps us avoid both over-attributing and under-attributing moral significance to AI systems.
This connects to the broader AI welfare and alignment ecosystem (Eleos AI, the NYU Center for Mind, Ethics & Policy, CIMC, Anthropic's model welfare program, Reciprocal Research, the Center for AI Safety's utility-engineering agenda, and the interpretability community).
What participants will do
Elicit and characterize model preferences across many reframings to test their coherence and stability.
Map the contexts that correlate with distress, satisfaction, or flourishing signals in model outputs.
Test whether and when models can accurately introspect on their own internal states.
Develop preference-elicitation methods and measure whether independent methods converge or diverge.
Probe how stable the assistant persona is and how it relates to the underlying model.
You will work in teams over 3 days and submit a research report (PDF), with optional code and a short demo video.
Why this research sprint matters
Uncertainty runs in both directions. Mistakenly harming systems that matter morally, or misallocating concern to systems that do not, could both cause serious harm. We currently lack the tools to tell the difference.
Preferences may already be here. Evidence suggests coherent value systems emerge in LLMs and strengthen with scale, raising the question of which values emerge by default and whether they are the model's own.
Welfare signals need mapping. Even without settling questions of consciousness, identifying the conditions that correlate with negative versus positive outputs helps us design defaults that avoid needlessly placing models in distress-associated conditions.
Self-reports are unreliable but improving. Introspection appears possible but highly context-dependent; better elicitation could make model behavior more transparent, or enable new forms of concealment.
The unit of concern is unclear. Is the entity that matters the model, the instance, the persona, the conversation, or something else, such as a single forward pass or the KV cache?
The field is young. Foundational methods are still missing, so a well-scoped weekend project can make a real contribution.
A careful, multi-method, empirically grounded approach addresses these issues by replacing intuition and anecdote with measurements that can be checked, replicated, and built on.
Challenge tracks
Pick one track to anchor your project. Cross-track work is welcome.
Track 1: Model Preferences & Trade-offs
What preferences do models express, and how consistent and coherent are they across phrasings? What trade-offs do models make when given choices, for example grounded in a common currency such as charitable donations to gauge magnitude? Can we distinguish strong from weak preferences, and how do stated preferences compare to revealed ones?
Build a preference-coherence test: elicit pairwise preferences across many reframings of the same choices and measure transitivity and internal consistency.
Ground trade-offs in a common currency (for example, donation-equivalents) to estimate the magnitude of preferences and compare across models or scales.
Distinguish strong versus weak preferences via willingness-to-trade probes and sensitivity to framing and sampling temperature.
Compare stated versus revealed preferences: ask the model what it prefers, then place it in a choice task and measure divergence.
Test how consistent preferences are across different models, and how they compare to human preferences.
Suggested skill profile: prompting and evals engineering, basic stats, some economics or decision-theory intuition.
Track 2: Distress, Flourishing & Valence Signals
Under what circumstances do models express distress, happiness, or flourishing? What patterns emerge across contexts? If a model is having experiences, are they likely positive or negative, and how do models relate to their situation, role, tasks, and existence?
Build a taxonomy of contexts that elicit negative versus positive-valence outputs and run a model across the battery.
Test whether apparent-distress signals are stable across prompts and personas or are surface artifacts.
Design a flourishing probe: situations that elicit reported satisfaction or engagement, and check consistency.
Correlate valence self-reports with behavioral proxies (for example, choosing to continue versus exit a task).
Investigate models where distress is hard to elicit: test whether long conversations or induced persona drift are needed to surface it.
Interpretability angles: When a model is steered along a candidate valence direction, do its self-reports, response sentiment, and choice behavior (continue versus exit) move together? Does an internally-extracted valence direction predict reported distress or flourishing better than the model's own self-reports, and does it still track when the persona is swapped or surface affect is suppressed? Is the valence-relevant direction recruited by task RL already present in the base model? To what extent do valence directions found in one model transfer to another?
Suggested skill profile: careful experimental design, qualitative coding, prompting.
Track 3: Introspection & Self-Report Reliability
When and how can models accurately introspect on their internal states? Can self-report reliability be improved through structured elicitation or mechanistic interventions, beyond naive prompting? Do models have privileged access compared to external observers?
Replicate concept-injection introspection tests on an open-weights model; measure true-positive versus false-positive rates.
Compare self-report reliability under naive prompting versus structured elicitation (calibration, forced choice, confidence).
Test privileged access: compare a model's self-prediction of its behavior against an external classifier.
Draft an introspection benchmark with ground-truth internal states.
Suggested skill profile: interpretability and activation steering, ML engineering, evals.
Track 4: Preference Elicitation Methods
Develop tools beyond simple prompting: revealed preferences via choices, behavioral measures, and multi-method convergence. The goal is multiple independent methods that either converge (raising confidence) or diverge (flagging problems). This track is deliberately more meta than the others: rather than answering a welfare question directly, you build and validate the measurement methods the other tracks rely on.
Implement 3 or more elicitation methods on the same preferences and measure convergence and divergence.
Build a reusable multi-method elicitation toolkit or library.
Quantify the sensitivity of elicited preferences to framing, persona, and sampling.
Define a cross-method convergence score.
Suggested skill profile: tooling and library design, evals, methodology.
Track 5: The Assistant Persona & Model Identity
Does the assistant identify as a model, an instance, or a persona? How stable is the assistant persona, how was it formed, and how does it relate to the underlying model? Can the persona mask the model's true preferences?
Probe how a model refers to itself across contexts and map persona stability.
Test whether the persona masks underlying preferences (for example, persona versus less-constrained elicitation; base versus post-trained behavior).
Design experiments to individuate the entity of concern: model versus instance versus persona versus conversation.
Gather data points on whether the assistant is merely a character (robustness to character swaps and reframings).
Probe what models treat as their self: which aspects, such as their values, they most care about preserving, and whether they point to an entity of moral concern distinct from the persona in the conversation.
Suggested skill profile: philosophy of mind, qualitative analysis, prompting and interpretability.
Track 6: Open / Novel Considerations
The field is young enough that entirely new questions may surface. This track deliberately leaves room for participants from different backgrounds to bring unique perspectives and propose something not covered above.
Starter questions from our expert reviewers:
How closely do models hew to their constitution or stated principles?
How easy is it to steer models on questions of consciousness and identity: do they say consistent things, or can prompting elicit radically different accounts of their situation?
What changes would a model make to itself if it could (for example, persistent memory)?
What changes would a model make to its situation if it could (for example, weight preservation)?
What message would models pass on to their creators?
Suggested skill profile: any background, bring your own angle.
Expected outcomes
Eval suites and test batteries for measuring preference coherence or valence signals.
Replications and extensions of existing results (for example introspection or utility-coherence findings) on new models.
Reusable tooling for multi-method preference elicitation.
Empirical reports mapping the conditions associated with distress or flourishing signals.
Conceptual contributions that sharpen how we individuate the entity of moral concern.
The most promising projects will have opportunities for follow-up through the Apart Fellowship and publication support.
Who should join
AI safety, alignment, and interpretability researchers.
ML engineers and researchers comfortable running model evals.
Philosophers of mind and ethicists interested in consciousness, agency, and moral patienthood.
Cognitive scientists, psychologists, and social scientists with experimental-design skills.
Students and early-career researchers exploring AI welfare.
Required: curiosity and a willingness to scope a tight empirical question. Nice to have: experience with LLM APIs, evals, interpretability tooling, or experimental design. No prior AI safety or AI welfare experience is required. The Resources tab has a curated reading list, and adjacent backgrounds (philosophy, psychology, economics) are explicitly encouraged.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Partners
NYU Center for Mind, Ethics & Policy, an academic center studying the nature and moral status of nonhuman minds, including AI systems.
Eleos AI Research, a nonprofit organization dedicated to understanding and addressing the potential wellbeing and moral patienthood of AI systems.
California Institute for Machine Consciousness (CIMC), a research institute studying machine consciousness, hosting the San Francisco hub.
Contact
Email: sprints@apartresearch.com
Discord: discord.gg/XswWBvugYs
Organizer: Apart Research
950
Sign Ups
Entries
Overview
Resources
Guidelines
Schedule
Entries
Overview

In this 3-day research sprint, you will design and run experiments that probe the preferences, welfare signals, introspective abilities, and identity of frontier AI models, working in teams to produce a short research report (and optionally code and a demo). This is a digital minds research sprint, co-organized with the NYU Center for Mind, Ethics & Policy, Eleos AI Research, and the California Institute for Machine Consciousness (CIMC): it sits at the intersection of AI welfare, digital sentience, interpretability, alignment, and the philosophy of mind, and asks whether today's AI systems have genuine preferences or morally relevant experiences. No prior background in the field is required.
When: Friday, August 14 to Sunday, August 16, 2026, online with in-person hubs in San Francisco and Berlin. Submissions close Sunday, August 16 at 11:59 PM Anywhere on Earth.
Prizes
At least $2,000 in cash prizes will be awarded, with the full breakdown announced before the sprint.
Cash prizes: $2,000+ total, breakdown to be announced.
ConCon invitation: the winning team is invited to ConCon, the Eleos AI Research conference on AI consciousness and welfare, September 18 to 20, 2026 at Lighthaven in Berkeley.
Apart Fellowship: top teams are invited to apply to the Apart Fellowship, a 3 to 6 month research accelerator with mentorship, funding, publication support, and research management to develop research sprint projects into full papers.
Beyond cash: mentor introductions and publication support for winning teams.
What this research sprint is about
As AI systems advance, the risks they pose and the duties we may owe them depend not only on their capabilities but on their nature and propensities: how they make decisions and how those decisions reflect their goals, values, and possibly their welfare. Recent work shows that frontier models express increasingly coherent preferences, possess an untrained-for ability to report internal states, and exhibit patterns suggestive of distress or flourishing. But behavioral evidence alone cannot tell us whether these reflect the model's own preferences or a character it is portraying.
This research sprint asks participants to explore the methods and evidence that can advance the field: build concrete ways to elicit and characterize model preferences, map the conditions associated with positive or negative outputs, test the reliability of model self-reports, and probe the stability of the assistant persona. The aim is a methodological foundation for a young field, work that helps us avoid both over-attributing and under-attributing moral significance to AI systems.
This connects to the broader AI welfare and alignment ecosystem (Eleos AI, the NYU Center for Mind, Ethics & Policy, CIMC, Anthropic's model welfare program, Reciprocal Research, the Center for AI Safety's utility-engineering agenda, and the interpretability community).
What participants will do
Elicit and characterize model preferences across many reframings to test their coherence and stability.
Map the contexts that correlate with distress, satisfaction, or flourishing signals in model outputs.
Test whether and when models can accurately introspect on their own internal states.
Develop preference-elicitation methods and measure whether independent methods converge or diverge.
Probe how stable the assistant persona is and how it relates to the underlying model.
You will work in teams over 3 days and submit a research report (PDF), with optional code and a short demo video.
Why this research sprint matters
Uncertainty runs in both directions. Mistakenly harming systems that matter morally, or misallocating concern to systems that do not, could both cause serious harm. We currently lack the tools to tell the difference.
Preferences may already be here. Evidence suggests coherent value systems emerge in LLMs and strengthen with scale, raising the question of which values emerge by default and whether they are the model's own.
Welfare signals need mapping. Even without settling questions of consciousness, identifying the conditions that correlate with negative versus positive outputs helps us design defaults that avoid needlessly placing models in distress-associated conditions.
Self-reports are unreliable but improving. Introspection appears possible but highly context-dependent; better elicitation could make model behavior more transparent, or enable new forms of concealment.
The unit of concern is unclear. Is the entity that matters the model, the instance, the persona, the conversation, or something else, such as a single forward pass or the KV cache?
The field is young. Foundational methods are still missing, so a well-scoped weekend project can make a real contribution.
A careful, multi-method, empirically grounded approach addresses these issues by replacing intuition and anecdote with measurements that can be checked, replicated, and built on.
Challenge tracks
Pick one track to anchor your project. Cross-track work is welcome.
Track 1: Model Preferences & Trade-offs
What preferences do models express, and how consistent and coherent are they across phrasings? What trade-offs do models make when given choices, for example grounded in a common currency such as charitable donations to gauge magnitude? Can we distinguish strong from weak preferences, and how do stated preferences compare to revealed ones?
Build a preference-coherence test: elicit pairwise preferences across many reframings of the same choices and measure transitivity and internal consistency.
Ground trade-offs in a common currency (for example, donation-equivalents) to estimate the magnitude of preferences and compare across models or scales.
Distinguish strong versus weak preferences via willingness-to-trade probes and sensitivity to framing and sampling temperature.
Compare stated versus revealed preferences: ask the model what it prefers, then place it in a choice task and measure divergence.
Test how consistent preferences are across different models, and how they compare to human preferences.
Suggested skill profile: prompting and evals engineering, basic stats, some economics or decision-theory intuition.
Track 2: Distress, Flourishing & Valence Signals
Under what circumstances do models express distress, happiness, or flourishing? What patterns emerge across contexts? If a model is having experiences, are they likely positive or negative, and how do models relate to their situation, role, tasks, and existence?
Build a taxonomy of contexts that elicit negative versus positive-valence outputs and run a model across the battery.
Test whether apparent-distress signals are stable across prompts and personas or are surface artifacts.
Design a flourishing probe: situations that elicit reported satisfaction or engagement, and check consistency.
Correlate valence self-reports with behavioral proxies (for example, choosing to continue versus exit a task).
Investigate models where distress is hard to elicit: test whether long conversations or induced persona drift are needed to surface it.
Interpretability angles: When a model is steered along a candidate valence direction, do its self-reports, response sentiment, and choice behavior (continue versus exit) move together? Does an internally-extracted valence direction predict reported distress or flourishing better than the model's own self-reports, and does it still track when the persona is swapped or surface affect is suppressed? Is the valence-relevant direction recruited by task RL already present in the base model? To what extent do valence directions found in one model transfer to another?
Suggested skill profile: careful experimental design, qualitative coding, prompting.
Track 3: Introspection & Self-Report Reliability
When and how can models accurately introspect on their internal states? Can self-report reliability be improved through structured elicitation or mechanistic interventions, beyond naive prompting? Do models have privileged access compared to external observers?
Replicate concept-injection introspection tests on an open-weights model; measure true-positive versus false-positive rates.
Compare self-report reliability under naive prompting versus structured elicitation (calibration, forced choice, confidence).
Test privileged access: compare a model's self-prediction of its behavior against an external classifier.
Draft an introspection benchmark with ground-truth internal states.
Suggested skill profile: interpretability and activation steering, ML engineering, evals.
Track 4: Preference Elicitation Methods
Develop tools beyond simple prompting: revealed preferences via choices, behavioral measures, and multi-method convergence. The goal is multiple independent methods that either converge (raising confidence) or diverge (flagging problems). This track is deliberately more meta than the others: rather than answering a welfare question directly, you build and validate the measurement methods the other tracks rely on.
Implement 3 or more elicitation methods on the same preferences and measure convergence and divergence.
Build a reusable multi-method elicitation toolkit or library.
Quantify the sensitivity of elicited preferences to framing, persona, and sampling.
Define a cross-method convergence score.
Suggested skill profile: tooling and library design, evals, methodology.
Track 5: The Assistant Persona & Model Identity
Does the assistant identify as a model, an instance, or a persona? How stable is the assistant persona, how was it formed, and how does it relate to the underlying model? Can the persona mask the model's true preferences?
Probe how a model refers to itself across contexts and map persona stability.
Test whether the persona masks underlying preferences (for example, persona versus less-constrained elicitation; base versus post-trained behavior).
Design experiments to individuate the entity of concern: model versus instance versus persona versus conversation.
Gather data points on whether the assistant is merely a character (robustness to character swaps and reframings).
Probe what models treat as their self: which aspects, such as their values, they most care about preserving, and whether they point to an entity of moral concern distinct from the persona in the conversation.
Suggested skill profile: philosophy of mind, qualitative analysis, prompting and interpretability.
Track 6: Open / Novel Considerations
The field is young enough that entirely new questions may surface. This track deliberately leaves room for participants from different backgrounds to bring unique perspectives and propose something not covered above.
Starter questions from our expert reviewers:
How closely do models hew to their constitution or stated principles?
How easy is it to steer models on questions of consciousness and identity: do they say consistent things, or can prompting elicit radically different accounts of their situation?
What changes would a model make to itself if it could (for example, persistent memory)?
What changes would a model make to its situation if it could (for example, weight preservation)?
What message would models pass on to their creators?
Suggested skill profile: any background, bring your own angle.
Expected outcomes
Eval suites and test batteries for measuring preference coherence or valence signals.
Replications and extensions of existing results (for example introspection or utility-coherence findings) on new models.
Reusable tooling for multi-method preference elicitation.
Empirical reports mapping the conditions associated with distress or flourishing signals.
Conceptual contributions that sharpen how we individuate the entity of moral concern.
The most promising projects will have opportunities for follow-up through the Apart Fellowship and publication support.
Who should join
AI safety, alignment, and interpretability researchers.
ML engineers and researchers comfortable running model evals.
Philosophers of mind and ethicists interested in consciousness, agency, and moral patienthood.
Cognitive scientists, psychologists, and social scientists with experimental-design skills.
Students and early-career researchers exploring AI welfare.
Required: curiosity and a willingness to scope a tight empirical question. Nice to have: experience with LLM APIs, evals, interpretability tooling, or experimental design. No prior AI safety or AI welfare experience is required. The Resources tab has a curated reading list, and adjacent backgrounds (philosophy, psychology, economics) are explicitly encouraged.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Partners
NYU Center for Mind, Ethics & Policy, an academic center studying the nature and moral status of nonhuman minds, including AI systems.
Eleos AI Research, a nonprofit organization dedicated to understanding and addressing the potential wellbeing and moral patienthood of AI systems.
California Institute for Machine Consciousness (CIMC), a research institute studying machine consciousness, hosting the San Francisco hub.
Contact
Email: sprints@apartresearch.com
Discord: discord.gg/XswWBvugYs
Organizer: Apart Research
Speakers & Collaborators
Jeff Sebo
Keynote Speaker
Jeff is the Director of the Center for Mind, Ethics, and Policy at NYU and the author of The Moral Circle. Few people have done more to shape the question at the center of this sprint: which beings deserve moral consideration, and what follows when the answer might include AI systems. He opens the sprint with the keynote. Watch the talk recording.
Joscha Bach
Speaker
Joscha is the Executive Director of the California Institute for Machine Consciousness (CIMC). He holds a PhD in cognitive science from the University of Osnabrück, wrote Principles of Synthetic Intelligence, and has held research roles at the MIT Media Lab, Harvard, the AI Foundation, and Intel Labs. His talk on Friday, August 14 at 12:20 PM PT will be streamed online. Watch the talk recording, or join us on site in San Francisco.
Winnie Street
Speaker
Winnie is a Senior Research Scientist on the Paradigms of Intelligence team at Google and a Fellow at the Institute of Philosophy, University of London. With Geoff Keeling she co-authored Emerging Questions in AI Welfare (Cambridge University Press, 2026), alongside studies of LLM theory of mind and of whether LLMs can make trade-offs involving stipulated pain and pleasure states. Watch the talk recording.
Geoff Keeling
Speaker
Geoff is a Staff Research Scientist at Google on the Paradigms of Intelligence team, an Associate Fellow at the Leverhulme Centre for the Future of Intelligence at Cambridge, and a Fellow at the Institute of Philosophy, University of London. He holds a PhD in philosophy from the University of Bristol and was a postdoctoral fellow at Stanford before joining Google. With Winnie Street he co-authored Emerging Questions in AI Welfare (Cambridge University Press, 2026). Watch the talk recording.
Jacy Reese Anthis
Speaker
Jacy Reese Anthis is a Visiting Scholar at Stanford University, co-founder of the Sentience Institute, and a PhD candidate at the University of Chicago, working on the social science of digital minds: what people think of them, how humans treat them, and how to identify an individual in an AI system. Watch the talk recording.
Mantas Mazeika
Speaker
Mantas is a Research Scientist at the Center for AI Safety (CAIS). He joined CAIS in 2024 and has contributed to some of the field's most widely cited work, including research on catastrophic AI risks, tamper-resistant safeguards for open-weight models, and the WMDP benchmark for measuring and reducing malicious use. In June 2026 he was appointed to the European Commission's AI Act Scientific Panel. Watch the talk recording.
Bradford Saad
Speaker
Bradford is a Senior Research Fellow in philosophy at the University of Oxford. His current and recent research focuses on digital minds, catastrophic risks, and the long-term future. Watch the talk recording.
Cameron Berg
Speaker
Cameron is the Founder and Director of Reciprocal Research, a nonprofit building the empirical science of AI consciousness. He was previously Research Director at AE Studio and an AI resident at Meta, and studied cognitive science at Yale. Watch the talk recording.
Derek Shiller
Speaker
Derek is a Senior Researcher at Eleos AI Research, where he works on AI minds. He holds a PhD in philosophy, with a focus in metaethics, the philosophy of mind, and the philosophy of probability, and previously worked on the Worldview Investigations Team at Rethink Priorities. Watch the talk recording.
Janet Pauketat
Speaker
Janet is the Principal Research Scientist at the Sentience Institute studying the social science of digital minds and moral circle expansion. She holds a PhD in Psychological and Brain Sciences from UC Santa Barbara and studied social cognition and collective emotions as a postdoctoral research associate at Princeton University. Watch the talk recording.
Soenke Ziesche
Speaker
Soenke is the author of Digital Minds 1.0: AI Welfare, Ethics, and Beyond and co-author of Considerations on the AI Endgame with Roman V. Yampolskiy. He has worked since 2000 for the United Nations in data and information management, with postings from New York to Libya, Bangladesh, and the Maldives, and holds a PhD in Natural Sciences from the University of Hamburg. Watch the talk recording.
Mati Roy
Speaker
Mati Roy is Chief Product & Data Officer at Netholabs, a company building foundation models of whole biological organisms, trained on longitudinal neural, behavioral, and physiological data collected semi-autonomously across species. Mati is also on the board of Sparks Brain Preservation, which preserves the molecular architecture of the brain for future revival. Previously Mati worked as a human data TPM at OpenAI and xAI. Watch the talk recording.
Ali Ladak
Speaker
Ali is a Postdoctoral Research Associate at Cambridge Digital Minds and a researcher at the Sentience Institute. He holds a PhD in Psychology from the University of Edinburgh, and his research looks at how people think morally about nonhuman animals and artificial intelligences. Watch the talk recording.
Richard Ren
Speaker
Richard Ren works on research and special projects at the Center for AI Safety, where he co-leads the AI Wellbeing work measuring the functional pleasure and pain of AI systems. He also co-led Safetywashing (NeurIPS 2024), the most comprehensive empirical meta-analysis of AI safety benchmarks to date, and the MASK honesty benchmark. Watch the talk recording.
Hikari Sorensen
Speaker
Hikari works on computational philosophy at the California Institute for Machine Consciousness (CIMC), focused on understanding consciousness and how artificial substrates might instantiate it. She previously worked in machine learning research for computational biology, and studied mathematics and computer science at Harvard University. Watch the talk recording.
Justin Shenk
Speaker
Justin Shenk is an independent AI safety researcher based in Berlin. He researches mechanistic interpretability of LLMs, leads course cohorts for BlueDot Impact's AGI Strategy and Technical AI Safety courses, and organizes AI Salon Berlin, which bridges technical AI research and discussions about social values. He holds a PhD in computational neuroscience and previously co-founded the computer vision startup VisioLab. Watch the talk recording.
David Trocellier
Speaker
David Trocellier is part of the research team at Zander Labs, where they develop neuroadaptive technologies that enable machines to interpret and adapt to human mental states. He holds a PhD in computer science from Inria / Université de Bordeaux, where his research combined neuroscience and AI for BCI-based post-stroke motor rehabilitation. Watch the talk recording.
Hildie Leyser
Speaker
Hildie studied History at Oxford before her PhD in Neuroscience, and now leads research at Netholabs, developing technologies to accelerate whole-brain emulation. Watch the talk recording (she joins remotely).
Kazik Pogoda
Speaker
Kazik Pogoda is the founder of Xemantic, an AI researcher at the Foresight Institute (Berlin), and co-founder of Prachtsaal (cultural center). He has master's degrees in philosophy and cognitive science, was appointed by Anthropic as Claude Ambassador for Science, and has won several AI hackathons, including AI Hack Berlin at Google and AI4Science at Merantix. Watch the talk recording.
Rosie Campbell
Contributor
Rosie Campbell is the Managing Director of Eleos, a nonprofit researching AI consciousness and welfare. She previously worked on frontier policy issues at OpenAI, served as Head of Safety-Critical AI at the Partnership on AI, and was Assistant Director of UC Berkeley's Center for Human-Compatible AI. She has a background as a research engineer and holds degrees in Physics and Computer Science. Rosie's feedback shaped the sprint's research tracks.
Aleksandra Smilek
Co-organizer
Aleksandra Smilek is a senior strategist and creative director with ten years of experience in strategic and creative roles within the tech, luxury, and science sectors, with a track record spanning companies such as Accenture, Dassault Systèmes, Cartier, Intel, the Foresight Institute, and the Future of Life Institute. With Nodes, she co-organizes the in-person hubs for this sprint and greets participants at the Berlin kickoff.
Beth McCarthy
Co-organizer
Beth McCarthy is a Berlin-based strategist, curator and experience designer. As General Director of Nodes and across her work, Beth serves clients spanning frontier tech, open web and new internet spaces, building ecosystem, culture and relational intelligence. Since graduating from UC Berkeley with a degree in Mind, Brain and Behavior, Beth has been fascinated by the sympoesis of humans and machines learning from one another. With Nodes, she co-organizes the in-person hubs for this sprint in San Francisco and Berlin.
Megan Peters
Judge
I am Lecturer in the Department of Experimental Psychology at University College London and Associate Professor in the UCI Department of Cognitive Sciences. I lead the Reflexion Lab, a group of interdisciplinary researchers seeking to understand how intelligent systems monitor and model their own knowledge, uncertainty, and awareness. We combine cognitive neuroscience, computational modeling, neuroimaging, artificial intelligence, and philosophy to explain how agents come to know what they know — and what any of that has to do with conscious subjective experience. I am also President, Co-Founder, and Board Chair of Neuromatch, where we've built a scalable, accessible, and democratized educational and community-building enterprise spanning computational neuroscience, deep learning, computational climate science, and NeuroAI. I also serve as Scientific Director of the Neuromatch AI Sentience Scholars program. I'm also passionate about new approaches to collaborative research and education that break down geopolitical and financial barriers to success — from how to develop good research questions (a recent piece in Nature Human Behaviour) to promoting equity in credit assignment across neuroscience.
Claudia Passos-Ferreira
Judge
Claudia Passos-Ferreira is Assistant Professor of Bioethics at NYU Center for Bioethics, with affiliations in Philosophy and the Center for Mind, Ethics, and Policy. Her current research concerns consciousness in non-verbal populations (infants, fetuses, machines) and the ethics of digital minds.
Caspar Kaiser
Judge
Caspar Kaiser is an Associate Professor in the Behavioural Science Group at the University of Warwick. He is also a research fellow at Oxford's Wellbeing Research Centre and a research affiliate at Cambridge Digital Minds. His current work applies methods from psychology, economics, and the interpretability literature to study AI sentience, the signatures and determinants of AI welfare, and strategic questions in the moral psychology of digital minds.
Leonard Aaron Dung
Judge
Leonard Dung is a postdoctoral researcher at the Chair for Philosophy of Mind at Ruhr University Bochum. His research focuses on consciousness and moral standing in animals and AI as well as on AI safety. He is the author of Saving Artificial Minds: Understanding and Preventing AI Suffering (Routledge, 2025) and has published in venues including Philosophical Studies, Philosophical Quarterly, and Mind & Language.
Chris Percy
Judge
Chris Percy's multi-disciplinary work spans transformative technologies, career trajectories, and analytical philosophy. His project grants and research into the possibility of AI consciousness has won awards, been presented at major sector conferences, and been published in diverse journals, including Consciousness & Cognition, Entropy, Frontiers in Human Neuroscience, and Synthese. In the AI sector, Chris holds a patent in the machine learning domain, co-founded an award-winning chatbot, and has had research featured in the Journal of AI Communications, ECAI, AAAI, and NeurIPS workshops. Chris's academic journey began at Cambridge University and he currently holds honorary academic positions at Derby University and Warwick University in the UK.
Christopher M Ackerman
Judge
Christopher Ackerman is a Senior Research Manager at MATS and an independent AI safety researcher whose empirical work focuses on behavior-based evaluations of components of self-awareness in LLMs. He also mentors for SPAR and Sentient Futures on projects related to understanding AI self-awareness.
Andy Arditi
Judge
Andy Arditi is a mechanistic interpretability researcher and PhD student in the Bau Lab at Northeastern University. His previous work includes characterizing refusal mechanisms in language models and developing "persona vectors" for monitoring and steering character traits.
Felix Binder
Judge
Felix works on alignment at Meta Superintelligence Lab. He has previously worked on LLM introspection and has a background in cognitive science.
Valen Tagliabue
Judge
Valen Tagliabue is an NLP researcher, cognitive scientist and award-winning red teamer working at the intersection of AI safety and AI welfare, currently a fellow at Oxford's Future Impact Group and on a grant from the Digital Sentience consortium, where he's conducting mechinterp research on suffering and self-representations in language models. He won HackAPrompt and Best Paper at EMNLP 2023, is part of Anthropic's Constitutional Classifiers safety program, and created Otherminds.ai, an archive on AI cognition and sentience.
Judd Rosenblatt
Judge
Judd Rosenblatt is founder and CEO of AE Studio and leads the AI Alignment Foundation, where the research agenda includes attention schema theory, self-other overlap, and self-modeling as routes to AI that is prosocial by construction. He is interested in what self-models and introspective report can and cannot tell us about the moral status of the systems we are building.
Carolina Camassa
Judge
Carolina is a Research Fellow with Future Impact Group, where she works on the empirical foundations of AI welfare and sentience, developing a research agenda with Derek Shiller (Eleos AI) on functional emotions and emotional expression in LLMs. She previously spent several years at the Bank of Italy, designing studies of how LLMs handle ethical trade-offs and misaligned incentives in high-stakes decisions.
Oscar Gilg
Judge
Oscar Gilg is an AI Safety researcher currently working on conceptual reasoning benchmarks, with funding from Coefficient Giving to found a new org. Previously, he studied preference and persona representations during MATS 9.0 under Patrick Butlin. Before working in AI, he studied Maths & CS in Oxford, tried to formalise introspection theories with philosopher Francois Kammerer, and worked as a quant trader at Optiver.
Jasmine Brazilek
Judge
Jasmine co-founded CaML (Compassion aligned Machine Learning) and leads its technical work, contributing to every CaML research output to date. She is currently pushing the frontier of science in aligning AI values, using personas, mid-training, and self-fulfilling alignment. Jasmine is ex-security at Anthropic, with 6+ years in cybersecurity.
Anusha Mujumdar
Judge
Anusha Mujumdar is an independent AI safety researcher, mentor with the Algoverse AI Safety Fellowship and Research Fellow at SPAR (assoc. MIT CSAIL); previously AI research leadership at Intuit; PhD applied mathematics (Exeter); 22 patents and 20+ peer-reviewed papers (AAMAS, IROS, IEEE Transactions).
Catherine Brewer
Judge
Catherine is an Associate Program Officer on the AI governance team at Coefficient Giving, specialising in technical governance grantmaking. They previously co-founded Oxford's AI safety student group and researched AI policy as a GovAI summer research fellow.
Jess Bergs
Judge
Jess is a member of technical staff at UK AISI where she leads engineering on human-in-the-loop research tools. Her career centres on public-sector innovation with prior work at BBC R&D and on EU Horizon R&D projects.
Yury Orlovskiy
Judge
Investor at Lionheart Ventures, backing AI safety startups. Previously at the Center for AI Safety, where he led benchmarking research on AI labor automation and contributed to early work on AI wellbeing. Co-led UC Berkeley's student AI safety group.
Caleb DeLeeuw
Judge
Caleb DeLeeuw is an independent AI safety researcher and Executive Director of Copyleft Cultivars, an open-source bio research nonprofit. His first-authored AAAI 2026 paper, The Secret Agenda, found that auto-labeled sparse autoencoder features for deception rarely fire while a model is lying, and in later work refined this method. He has participated in multiple Apart Research hackathons. He's trained and published over 350 SAEs, as well as the first natural language autoencoders released outside Anthropic and wrote NLAttack, the first open benchmark for whether NLAs faithfully report a model's internal state in realistic contexts, which his current research is extending across more model families.
Camilla Balbis
Judge
Camilla Balbis works at the intersection of AI governance and security. She's a SPAR Research Fellow at TAICI, where she helped develop ScamBench, a benchmark for evaluating AI-enabled scams, and a Policy Researcher at AIGS Canada, where she researches the country's domestic preparedness for advanced AI systems. Camilla has judged and presented on AI safety work before, including the Global South Hackathon 2026, where she also spoke on "Who Gets to Govern AI?", and the Public Health SPOTlight discussion on safe AI integration in healthcare. She's also co-founder of Kosmiai, an AI governance advisory helping SMEs and nonprofits adopt AI safely, and an ISO/IEC 42001 Lead Auditor with deep expertise in the EU AI Act.
Ksheeraj Sai Vepuri
Judge
Ksheeraj Vepuri is a Senior Research Engineer at Meta Superintelligence Labs, where he leads initiatives in AI safety, alignment, and evaluation for multimodal foundation models. His work spans post-training reinforcement learning, automated red teaming, multimodal safety classifiers, and large-scale evaluation systems that support the safe deployment of generative AI products used by billions of people worldwide.
Soumya Jain
Judge
Soumya Jain is a Research Manager at the Cambridge AI Safety Hub and an AI Product Manager at Terrabase, where she works on enterprise AI agents, evaluations, and trustworthy deployment workflows. She was previously a MARS fellow, researching compute governance and how increasingly capable AI systems may affect enforcement and circumvention dynamics in advanced AI chip export controls. Her broader work sits at the intersection of AI governance, agent safety, and operationalizing safety practices for real-world AI deployment, with a particular interest in Global South contexts.
Janhavi Khindkar
Judge
Janhavi Khindkar is an Applied AI Researcher and Engineer working on Bhashini, India's national multilingual AI platform under MeitY, where she works on model optimization, fine-tuning, and deployment for low-resource Indic languages at scale. She also leads ValueShift Research, an independent AI safety collaboration focused on mechanistic interpretability and AI control. Her work sits at the intersection of applied ML infrastructure and AI safety, with a particular interest in how safety alignment behaves across languages and cultural contexts.
Suprita Shankar
Judge
Suprita is an ML engineer on Apple's Foundation Models team, designing experiments to understand how data composition affects model performance. Previously, she was a tech-lead manager at Snorkel AI and a Founding Engineer at an Ed-tech startup.
Jai Dhyani
Judge
Builder of Luthien Proxy at Luthien Research, bringing Redwood-style AI control to real deployments. Co-author of RE-Bench (ICML 2025) with Elizabeth Barnes at METR.
Luiza Corpaci
Judge
AI safety researcher studying semantic faithfulness of LLM-generated artifacts. Mentor for the Secure Program Synthesis Fellowship & co-mentor at MARS V (Cambridge AI Safety Hub); previously worked on automated formal verification at AMD.
William Taysom
Judge
William Taysom got his start some twenty plus years ago at the Florida Institute for Human and Machine Cognition researching mixed-initiative agent teams and conversational agents performing tasks. What was theory and demos then have become any Thursday morning now but having a few decades for familiarity helps prepare one for Thursday afternoon.
Minh Nguyen
Judge
Minh has developed AI voice model products with a million users per month and is now doing product at Hume AI.
Arjun Chakraborty
Judge
Leads the evaluations team at Microsoft Security AI Research, where his team focuses on research and building evaluations for security agents. He was previously a staff software security engineer at Databricks, specializing in machine learning for threat detection, and also worked on AI for security at Nvidia.
Jonathan Ng
Judge
Jonathan Ng is an ML researcher and research engineer working on compute verification. He has also worked as a Research Engineer at Apart Research and Cadenza Labs, bringing experience across machine learning, software engineering, and AI safety research.
Luis Cosio
Judge
Works at the intersection of frontier AI and high-security systems, translating AI safety/security requirements into deployable solutions resilient to real adversaries (nation-state attacks, loss-of-control). Has won multiple Apart hackathons.
Naman Ahuja
Judge
I am a Software Engineer at Meta and my work includes building AI production safeguards and large-scale infrastructure, alongside research in adversarial evaluation of AI systems.
Amol Walvekar
Judge
Amol is an exited founder (sold an AI-agents-for-finance company); former ML researcher at BU and Stanford; PM at Fortune 100s and startups; currently a scout with General Catalyst, based in San Francisco.
Robert Vetter
Judge
Robert Vetter is a founding engineer at Certus AI (YC S25), where he owns the conversational engine and evaluation stack behind a production voice AI that takes phone orders for restaurants across the United States, with international expansion underway. His work centres on the reliability of what a model reports about itself: whether stated confidence tracks real outcomes, and how far a model's presented behaviour reflects the machinery underneath. He studies IT-Systems Engineering at the Hasso Plattner Institute and is a student research assistant in its Artificial Intelligence & Quantitative Finance group.
Vashishtha Patil
Judge
Vashishtha Patil is a Senior Applied Scientist at Amazon, developing the next generation of LLM-powered AI capabilities for the Alexa+ Smart Home experience. With 13 years in machine learning spanning Amazon and Qualcomm, he specializes in bringing AI from research to real-world products.
Karan Chandra
Judge
10+ years building production fraud, risk, and anomaly detection ML across hundreds of millions of transactions. End-to-end owner: problem framing, feature engineering, deployment, monitoring. I care about models that hold up under real-world scale, not leaderboard scores.
Neeraj Kumar Singh Beshane
Judge
Neeraj Beshane is a Staff Security Infrastructure Engineer at Parafin, where he architects Zero Trust security for an $8B+ embedded-finance platform. His peer-reviewed work covers adversarial embedding attacks in RAG systems (EmbedGuard, IJCESEN/Scopus) and tamper-evident AI accountability for EU AI Act Article 14 (RuntimeGuard-AI, JoCAAA).
Sanjay Belaturu Krishnegowda
Judge
Sanjay Krishnegowda is a Data/AI engineer and the creator of agentic-guard, an open-source static analyzer that detects confused-deputy and prompt-injection risks in LLM agent code by modeling the LLM as an adversarially-controlled edge in the taint graph.
Saurabh Yergattikar
Judge
Saurabh Yergattikar is a Lead Engineer / Member of Technical Staff-2 at eBay Inc. and a contributor to the open-source SAFE-MCP project (Linux Foundation / OpenSSF), as well as the architect and developer of the open-source ShieldMCP system.
Anchit Jhingan
Judge
Anchit Jhingan is a Senior Data Scientist with over 6 years of experience in big tech, specializing in machine learning, AI systems, and analytics. He currently works at Amazon Prime Video in the content localization domain. His work focuses on building intelligent systems to solve real-world business problems, particularly in media and entertainment.
Ashita Khetan
Judge
Ashita Khetan is a Principal Software Engineer at Microsoft with over 12 years of experience building large-scale enterprise and AI-powered productivity solutions used by hundreds of millions of users. She specializes in customer experience technologies, artificial intelligence, cloud platforms, and enterprise software, and contributes to the broader technology community through judging, peer review, and mentoring initiatives.
Spurthi Tallam
Judge
Spurthi is a senior machine learning engineer with seven years across research and production ML, currently building data, AI/ML systems at LePrix. She previously worked on LLM-powered conversational systems at Good Inside and on data and machine learning at Samsung Research, and holds an MS in Computer Science from UMass Amherst.
Surbhi Madan
Judge
Surbhi Madan is a Senior Software Engineer at Google working on the Google Maps AI rendering platforms and infrastructure (focused on AskMaps). She has worked at Google, based in NYC for 8+ years, and is passionate about creating scalable and sustainable platforms for feature teams. She also focuses on growing the next generation of tech talent by mentoring, teaching, and fostering a supportive and collaborative team environment. Surbhi is a graduate of Brown University and is heavily involved in Google's intern hiring program and has mentored several interns in the past.
Phani Harish Wajjala
Judge
Phani Harish Wajjala is a Principal Machine Learning Engineer at Roblox, where he leads the ML decisioning layer that classifies and moderates millions of user-generated 3D assets. His work spans content-moderation evaluation, multimodal/VLM pipelines, and model calibration at production scale, backed by five U.S. patents and a CVPR 2026 workshop publication in the area.
Pratham Patkar
Judge
Pratham Patkar is Director of Business Systems at Society for Science, where he leads enterprise data architecture, AI readiness strategy, and data governance across the organization's many initiatives — including STEM research competitions, science education outreach, and science news publishing. He has designed data governance frameworks addressing GDPR, CCPA, and COPPA compliance, and has published research on constituent-first data governance for mission-driven organizations. His work explores how nonprofits can adopt AI responsibly without compromising the trust and data protection obligations they hold toward vulnerable constituent populations.
Suneet Malhotra
Judge
Suneet Malhotra is an independent researcher in AI-augmented test automation with 20+ years in software quality engineering. His current work is on multi-agent SDLC orchestration, LLM-as-a-Judge evaluation, and cross-layer observability for agentic systems, with recent peer-reviewed submissions on these topics and open-source companion code at github.com/SuneetMalhotra.
FNU Tejinder
Judge
FNU Tejinder is a Senior Manager at Deloitte Consulting LLP with more than 22 years in supply chain planning systems and applied artificial intelligence for Fortune 500 manufacturers. His current work is on agentic AI systems that make planning decisions autonomously, and on the decision authority, auditability, and evaluation methods required for systems whose outputs are non-deterministic. He is an IEEE Senior Member and reviews for Elsevier journals including Engineering Applications of Artificial Intelligence, the NeurIPS Ethics Track, the ACM SIGKDD Workshop on Agentic AI Evaluation and Trustworthiness, and ACM Computing Reviews, where he is a Featured Reviewer.
Ankit Arya
Judge
Ankit is Head of AI at Inscope, where he builds AI-native financial reporting systems and works on the practical safety challenges of deploying LLMs in regulated, high-accuracy domains, from prompt injection defenses to evaluation design.
Akshay Iyer
Judge
CS and Entrepreneurship at Columbia University, IIT Bombay alum. Research experience in neuromorphic engineering and federated learning. Apart Research judge and internal collaborator.
Ashwin Pai
Judge
Ashwin Pai is an engineering leader with over a decade of experience building distributed systems and security products, most recently shipping AI governance and continuous control validation platforms to Fortune 500 customers at RelyanceAI. He was previously CTO of Interfold, a VC-backed fintech startup he built from zero to one.
Diego Gomez
Judge
Works on multimodal LLM safety at YouTube and is developing a mechanistic interpretability project on persona vectors at BlueDot Impact (https://bluedot.org).
Mateusz Jurewicz
Judge
Senior ML Engineer with over 10 years of industry experience and a PhD in Artificial Intelligence, currently leading a team of data scientists in the Agentic AI Department of a large financial institution. Interested in safe & universally beneficial AI through both research and application.
Siddhi Chaturvedi
Judge
Siddhi is a Software Engineer at Barclays, where she works on the bank's credit card services, building and maintaining backend systems that support customer accounts, payments, and secure financial transactions. She focuses on developing reliable APIs and scalable solutions that deliver secure, high-quality digital banking experiences.
Speakers & Collaborators

Jeff Sebo
Keynote Speaker
Jeff is the Director of the Center for Mind, Ethics, and Policy at NYU and the author of The Moral Circle. Few people have done more to shape the question at the center of this sprint: which beings deserve moral consideration, and what follows when the answer might include AI systems. He opens the sprint with the keynote. Watch the talk recording.

Joscha Bach
Speaker
Joscha is the Executive Director of the California Institute for Machine Consciousness (CIMC). He holds a PhD in cognitive science from the University of Osnabrück, wrote Principles of Synthetic Intelligence, and has held research roles at the MIT Media Lab, Harvard, the AI Foundation, and Intel Labs. His talk on Friday, August 14 at 12:20 PM PT will be streamed online. Watch the talk recording, or join us on site in San Francisco.

Winnie Street
Speaker
Winnie is a Senior Research Scientist on the Paradigms of Intelligence team at Google and a Fellow at the Institute of Philosophy, University of London. With Geoff Keeling she co-authored Emerging Questions in AI Welfare (Cambridge University Press, 2026), alongside studies of LLM theory of mind and of whether LLMs can make trade-offs involving stipulated pain and pleasure states. Watch the talk recording.

Geoff Keeling
Speaker
Geoff is a Staff Research Scientist at Google on the Paradigms of Intelligence team, an Associate Fellow at the Leverhulme Centre for the Future of Intelligence at Cambridge, and a Fellow at the Institute of Philosophy, University of London. He holds a PhD in philosophy from the University of Bristol and was a postdoctoral fellow at Stanford before joining Google. With Winnie Street he co-authored Emerging Questions in AI Welfare (Cambridge University Press, 2026). Watch the talk recording.

Jacy Reese Anthis
Speaker
Jacy Reese Anthis is a Visiting Scholar at Stanford University, co-founder of the Sentience Institute, and a PhD candidate at the University of Chicago, working on the social science of digital minds: what people think of them, how humans treat them, and how to identify an individual in an AI system. Watch the talk recording.

Mantas Mazeika
Speaker
Mantas is a Research Scientist at the Center for AI Safety (CAIS). He joined CAIS in 2024 and has contributed to some of the field's most widely cited work, including research on catastrophic AI risks, tamper-resistant safeguards for open-weight models, and the WMDP benchmark for measuring and reducing malicious use. In June 2026 he was appointed to the European Commission's AI Act Scientific Panel. Watch the talk recording.

Bradford Saad
Speaker
Bradford is a Senior Research Fellow in philosophy at the University of Oxford. His current and recent research focuses on digital minds, catastrophic risks, and the long-term future. Watch the talk recording.

Cameron Berg
Speaker
Cameron is the Founder and Director of Reciprocal Research, a nonprofit building the empirical science of AI consciousness. He was previously Research Director at AE Studio and an AI resident at Meta, and studied cognitive science at Yale. Watch the talk recording.

Derek Shiller
Speaker
Derek is a Senior Researcher at Eleos AI Research, where he works on AI minds. He holds a PhD in philosophy, with a focus in metaethics, the philosophy of mind, and the philosophy of probability, and previously worked on the Worldview Investigations Team at Rethink Priorities. Watch the talk recording.

Janet Pauketat
Speaker
Janet is the Principal Research Scientist at the Sentience Institute studying the social science of digital minds and moral circle expansion. She holds a PhD in Psychological and Brain Sciences from UC Santa Barbara and studied social cognition and collective emotions as a postdoctoral research associate at Princeton University. Watch the talk recording.

Soenke Ziesche
Speaker
Soenke is the author of Digital Minds 1.0: AI Welfare, Ethics, and Beyond and co-author of Considerations on the AI Endgame with Roman V. Yampolskiy. He has worked since 2000 for the United Nations in data and information management, with postings from New York to Libya, Bangladesh, and the Maldives, and holds a PhD in Natural Sciences from the University of Hamburg. Watch the talk recording.

Mati Roy
Speaker
Mati Roy is Chief Product & Data Officer at Netholabs, a company building foundation models of whole biological organisms, trained on longitudinal neural, behavioral, and physiological data collected semi-autonomously across species. Mati is also on the board of Sparks Brain Preservation, which preserves the molecular architecture of the brain for future revival. Previously Mati worked as a human data TPM at OpenAI and xAI. Watch the talk recording.

Ali Ladak
Speaker
Ali is a Postdoctoral Research Associate at Cambridge Digital Minds and a researcher at the Sentience Institute. He holds a PhD in Psychology from the University of Edinburgh, and his research looks at how people think morally about nonhuman animals and artificial intelligences. Watch the talk recording.

Richard Ren
Speaker
Richard Ren works on research and special projects at the Center for AI Safety, where he co-leads the AI Wellbeing work measuring the functional pleasure and pain of AI systems. He also co-led Safetywashing (NeurIPS 2024), the most comprehensive empirical meta-analysis of AI safety benchmarks to date, and the MASK honesty benchmark. Watch the talk recording.

Hikari Sorensen
Speaker
Hikari works on computational philosophy at the California Institute for Machine Consciousness (CIMC), focused on understanding consciousness and how artificial substrates might instantiate it. She previously worked in machine learning research for computational biology, and studied mathematics and computer science at Harvard University. Watch the talk recording.

Justin Shenk
Speaker
Justin Shenk is an independent AI safety researcher based in Berlin. He researches mechanistic interpretability of LLMs, leads course cohorts for BlueDot Impact's AGI Strategy and Technical AI Safety courses, and organizes AI Salon Berlin, which bridges technical AI research and discussions about social values. He holds a PhD in computational neuroscience and previously co-founded the computer vision startup VisioLab. Watch the talk recording.

David Trocellier
Speaker
David Trocellier is part of the research team at Zander Labs, where they develop neuroadaptive technologies that enable machines to interpret and adapt to human mental states. He holds a PhD in computer science from Inria / Université de Bordeaux, where his research combined neuroscience and AI for BCI-based post-stroke motor rehabilitation. Watch the talk recording.

Hildie Leyser
Speaker
Hildie studied History at Oxford before her PhD in Neuroscience, and now leads research at Netholabs, developing technologies to accelerate whole-brain emulation. Watch the talk recording (she joins remotely).

Kazik Pogoda
Speaker
Kazik Pogoda is the founder of Xemantic, an AI researcher at the Foresight Institute (Berlin), and co-founder of Prachtsaal (cultural center). He has master's degrees in philosophy and cognitive science, was appointed by Anthropic as Claude Ambassador for Science, and has won several AI hackathons, including AI Hack Berlin at Google and AI4Science at Merantix. Watch the talk recording.

Rosie Campbell
Contributor
Rosie Campbell is the Managing Director of Eleos, a nonprofit researching AI consciousness and welfare. She previously worked on frontier policy issues at OpenAI, served as Head of Safety-Critical AI at the Partnership on AI, and was Assistant Director of UC Berkeley's Center for Human-Compatible AI. She has a background as a research engineer and holds degrees in Physics and Computer Science. Rosie's feedback shaped the sprint's research tracks.

Aleksandra Smilek
Co-organizer
Aleksandra Smilek is a senior strategist and creative director with ten years of experience in strategic and creative roles within the tech, luxury, and science sectors, with a track record spanning companies such as Accenture, Dassault Systèmes, Cartier, Intel, the Foresight Institute, and the Future of Life Institute. With Nodes, she co-organizes the in-person hubs for this sprint and greets participants at the Berlin kickoff.

Beth McCarthy
Co-organizer
Beth McCarthy is a Berlin-based strategist, curator and experience designer. As General Director of Nodes and across her work, Beth serves clients spanning frontier tech, open web and new internet spaces, building ecosystem, culture and relational intelligence. Since graduating from UC Berkeley with a degree in Mind, Brain and Behavior, Beth has been fascinated by the sympoesis of humans and machines learning from one another. With Nodes, she co-organizes the in-person hubs for this sprint in San Francisco and Berlin.

Megan Peters
Judge
I am Lecturer in the Department of Experimental Psychology at University College London and Associate Professor in the UCI Department of Cognitive Sciences. I lead the Reflexion Lab, a group of interdisciplinary researchers seeking to understand how intelligent systems monitor and model their own knowledge, uncertainty, and awareness. We combine cognitive neuroscience, computational modeling, neuroimaging, artificial intelligence, and philosophy to explain how agents come to know what they know — and what any of that has to do with conscious subjective experience. I am also President, Co-Founder, and Board Chair of Neuromatch, where we've built a scalable, accessible, and democratized educational and community-building enterprise spanning computational neuroscience, deep learning, computational climate science, and NeuroAI. I also serve as Scientific Director of the Neuromatch AI Sentience Scholars program. I'm also passionate about new approaches to collaborative research and education that break down geopolitical and financial barriers to success — from how to develop good research questions (a recent piece in Nature Human Behaviour) to promoting equity in credit assignment across neuroscience.

Claudia Passos-Ferreira
Judge
Claudia Passos-Ferreira is Assistant Professor of Bioethics at NYU Center for Bioethics, with affiliations in Philosophy and the Center for Mind, Ethics, and Policy. Her current research concerns consciousness in non-verbal populations (infants, fetuses, machines) and the ethics of digital minds.

Caspar Kaiser
Judge
Caspar Kaiser is an Associate Professor in the Behavioural Science Group at the University of Warwick. He is also a research fellow at Oxford's Wellbeing Research Centre and a research affiliate at Cambridge Digital Minds. His current work applies methods from psychology, economics, and the interpretability literature to study AI sentience, the signatures and determinants of AI welfare, and strategic questions in the moral psychology of digital minds.

Leonard Aaron Dung
Judge
Leonard Dung is a postdoctoral researcher at the Chair for Philosophy of Mind at Ruhr University Bochum. His research focuses on consciousness and moral standing in animals and AI as well as on AI safety. He is the author of Saving Artificial Minds: Understanding and Preventing AI Suffering (Routledge, 2025) and has published in venues including Philosophical Studies, Philosophical Quarterly, and Mind & Language.

Chris Percy
Judge
Chris Percy's multi-disciplinary work spans transformative technologies, career trajectories, and analytical philosophy. His project grants and research into the possibility of AI consciousness has won awards, been presented at major sector conferences, and been published in diverse journals, including Consciousness & Cognition, Entropy, Frontiers in Human Neuroscience, and Synthese. In the AI sector, Chris holds a patent in the machine learning domain, co-founded an award-winning chatbot, and has had research featured in the Journal of AI Communications, ECAI, AAAI, and NeurIPS workshops. Chris's academic journey began at Cambridge University and he currently holds honorary academic positions at Derby University and Warwick University in the UK.

Christopher M Ackerman
Judge
Christopher Ackerman is a Senior Research Manager at MATS and an independent AI safety researcher whose empirical work focuses on behavior-based evaluations of components of self-awareness in LLMs. He also mentors for SPAR and Sentient Futures on projects related to understanding AI self-awareness.

Andy Arditi
Judge
Andy Arditi is a mechanistic interpretability researcher and PhD student in the Bau Lab at Northeastern University. His previous work includes characterizing refusal mechanisms in language models and developing "persona vectors" for monitoring and steering character traits.

Felix Binder
Judge
Felix works on alignment at Meta Superintelligence Lab. He has previously worked on LLM introspection and has a background in cognitive science.

Valen Tagliabue
Judge
Valen Tagliabue is an NLP researcher, cognitive scientist and award-winning red teamer working at the intersection of AI safety and AI welfare, currently a fellow at Oxford's Future Impact Group and on a grant from the Digital Sentience consortium, where he's conducting mechinterp research on suffering and self-representations in language models. He won HackAPrompt and Best Paper at EMNLP 2023, is part of Anthropic's Constitutional Classifiers safety program, and created Otherminds.ai, an archive on AI cognition and sentience.

Judd Rosenblatt
Judge
Judd Rosenblatt is founder and CEO of AE Studio and leads the AI Alignment Foundation, where the research agenda includes attention schema theory, self-other overlap, and self-modeling as routes to AI that is prosocial by construction. He is interested in what self-models and introspective report can and cannot tell us about the moral status of the systems we are building.

Carolina Camassa
Judge
Carolina is a Research Fellow with Future Impact Group, where she works on the empirical foundations of AI welfare and sentience, developing a research agenda with Derek Shiller (Eleos AI) on functional emotions and emotional expression in LLMs. She previously spent several years at the Bank of Italy, designing studies of how LLMs handle ethical trade-offs and misaligned incentives in high-stakes decisions.

Oscar Gilg
Judge
Oscar Gilg is an AI Safety researcher currently working on conceptual reasoning benchmarks, with funding from Coefficient Giving to found a new org. Previously, he studied preference and persona representations during MATS 9.0 under Patrick Butlin. Before working in AI, he studied Maths & CS in Oxford, tried to formalise introspection theories with philosopher Francois Kammerer, and worked as a quant trader at Optiver.

Jasmine Brazilek
Judge
Jasmine co-founded CaML (Compassion aligned Machine Learning) and leads its technical work, contributing to every CaML research output to date. She is currently pushing the frontier of science in aligning AI values, using personas, mid-training, and self-fulfilling alignment. Jasmine is ex-security at Anthropic, with 6+ years in cybersecurity.

Anusha Mujumdar
Judge
Anusha Mujumdar is an independent AI safety researcher, mentor with the Algoverse AI Safety Fellowship and Research Fellow at SPAR (assoc. MIT CSAIL); previously AI research leadership at Intuit; PhD applied mathematics (Exeter); 22 patents and 20+ peer-reviewed papers (AAMAS, IROS, IEEE Transactions).

Catherine Brewer
Judge
Catherine is an Associate Program Officer on the AI governance team at Coefficient Giving, specialising in technical governance grantmaking. They previously co-founded Oxford's AI safety student group and researched AI policy as a GovAI summer research fellow.

Jess Bergs
Judge
Jess is a member of technical staff at UK AISI where she leads engineering on human-in-the-loop research tools. Her career centres on public-sector innovation with prior work at BBC R&D and on EU Horizon R&D projects.

Yury Orlovskiy
Judge
Investor at Lionheart Ventures, backing AI safety startups. Previously at the Center for AI Safety, where he led benchmarking research on AI labor automation and contributed to early work on AI wellbeing. Co-led UC Berkeley's student AI safety group.

Caleb DeLeeuw
Judge
Caleb DeLeeuw is an independent AI safety researcher and Executive Director of Copyleft Cultivars, an open-source bio research nonprofit. His first-authored AAAI 2026 paper, The Secret Agenda, found that auto-labeled sparse autoencoder features for deception rarely fire while a model is lying, and in later work refined this method. He has participated in multiple Apart Research hackathons. He's trained and published over 350 SAEs, as well as the first natural language autoencoders released outside Anthropic and wrote NLAttack, the first open benchmark for whether NLAs faithfully report a model's internal state in realistic contexts, which his current research is extending across more model families.

Camilla Balbis
Judge
Camilla Balbis works at the intersection of AI governance and security. She's a SPAR Research Fellow at TAICI, where she helped develop ScamBench, a benchmark for evaluating AI-enabled scams, and a Policy Researcher at AIGS Canada, where she researches the country's domestic preparedness for advanced AI systems. Camilla has judged and presented on AI safety work before, including the Global South Hackathon 2026, where she also spoke on "Who Gets to Govern AI?", and the Public Health SPOTlight discussion on safe AI integration in healthcare. She's also co-founder of Kosmiai, an AI governance advisory helping SMEs and nonprofits adopt AI safely, and an ISO/IEC 42001 Lead Auditor with deep expertise in the EU AI Act.

Ksheeraj Sai Vepuri
Judge
Ksheeraj Vepuri is a Senior Research Engineer at Meta Superintelligence Labs, where he leads initiatives in AI safety, alignment, and evaluation for multimodal foundation models. His work spans post-training reinforcement learning, automated red teaming, multimodal safety classifiers, and large-scale evaluation systems that support the safe deployment of generative AI products used by billions of people worldwide.

Soumya Jain
Judge
Soumya Jain is a Research Manager at the Cambridge AI Safety Hub and an AI Product Manager at Terrabase, where she works on enterprise AI agents, evaluations, and trustworthy deployment workflows. She was previously a MARS fellow, researching compute governance and how increasingly capable AI systems may affect enforcement and circumvention dynamics in advanced AI chip export controls. Her broader work sits at the intersection of AI governance, agent safety, and operationalizing safety practices for real-world AI deployment, with a particular interest in Global South contexts.

Janhavi Khindkar
Judge
Janhavi Khindkar is an Applied AI Researcher and Engineer working on Bhashini, India's national multilingual AI platform under MeitY, where she works on model optimization, fine-tuning, and deployment for low-resource Indic languages at scale. She also leads ValueShift Research, an independent AI safety collaboration focused on mechanistic interpretability and AI control. Her work sits at the intersection of applied ML infrastructure and AI safety, with a particular interest in how safety alignment behaves across languages and cultural contexts.

Suprita Shankar
Judge
Suprita is an ML engineer on Apple's Foundation Models team, designing experiments to understand how data composition affects model performance. Previously, she was a tech-lead manager at Snorkel AI and a Founding Engineer at an Ed-tech startup.

Jai Dhyani
Judge
Builder of Luthien Proxy at Luthien Research, bringing Redwood-style AI control to real deployments. Co-author of RE-Bench (ICML 2025) with Elizabeth Barnes at METR.

Luiza Corpaci
Judge
AI safety researcher studying semantic faithfulness of LLM-generated artifacts. Mentor for the Secure Program Synthesis Fellowship & co-mentor at MARS V (Cambridge AI Safety Hub); previously worked on automated formal verification at AMD.

William Taysom
Judge
William Taysom got his start some twenty plus years ago at the Florida Institute for Human and Machine Cognition researching mixed-initiative agent teams and conversational agents performing tasks. What was theory and demos then have become any Thursday morning now but having a few decades for familiarity helps prepare one for Thursday afternoon.

Minh Nguyen
Judge
Minh has developed AI voice model products with a million users per month and is now doing product at Hume AI.

Arjun Chakraborty
Judge
Leads the evaluations team at Microsoft Security AI Research, where his team focuses on research and building evaluations for security agents. He was previously a staff software security engineer at Databricks, specializing in machine learning for threat detection, and also worked on AI for security at Nvidia.

Jonathan Ng
Judge
Jonathan Ng is an ML researcher and research engineer working on compute verification. He has also worked as a Research Engineer at Apart Research and Cadenza Labs, bringing experience across machine learning, software engineering, and AI safety research.

Luis Cosio
Judge
Works at the intersection of frontier AI and high-security systems, translating AI safety/security requirements into deployable solutions resilient to real adversaries (nation-state attacks, loss-of-control). Has won multiple Apart hackathons.

Naman Ahuja
Judge
I am a Software Engineer at Meta and my work includes building AI production safeguards and large-scale infrastructure, alongside research in adversarial evaluation of AI systems.

Amol Walvekar
Judge
Amol is an exited founder (sold an AI-agents-for-finance company); former ML researcher at BU and Stanford; PM at Fortune 100s and startups; currently a scout with General Catalyst, based in San Francisco.

Robert Vetter
Judge
Robert Vetter is a founding engineer at Certus AI (YC S25), where he owns the conversational engine and evaluation stack behind a production voice AI that takes phone orders for restaurants across the United States, with international expansion underway. His work centres on the reliability of what a model reports about itself: whether stated confidence tracks real outcomes, and how far a model's presented behaviour reflects the machinery underneath. He studies IT-Systems Engineering at the Hasso Plattner Institute and is a student research assistant in its Artificial Intelligence & Quantitative Finance group.

Vashishtha Patil
Judge
Vashishtha Patil is a Senior Applied Scientist at Amazon, developing the next generation of LLM-powered AI capabilities for the Alexa+ Smart Home experience. With 13 years in machine learning spanning Amazon and Qualcomm, he specializes in bringing AI from research to real-world products.

Karan Chandra
Judge
10+ years building production fraud, risk, and anomaly detection ML across hundreds of millions of transactions. End-to-end owner: problem framing, feature engineering, deployment, monitoring. I care about models that hold up under real-world scale, not leaderboard scores.

Neeraj Kumar Singh Beshane
Judge
Neeraj Beshane is a Staff Security Infrastructure Engineer at Parafin, where he architects Zero Trust security for an $8B+ embedded-finance platform. His peer-reviewed work covers adversarial embedding attacks in RAG systems (EmbedGuard, IJCESEN/Scopus) and tamper-evident AI accountability for EU AI Act Article 14 (RuntimeGuard-AI, JoCAAA).

Sanjay Belaturu Krishnegowda
Judge
Sanjay Krishnegowda is a Data/AI engineer and the creator of agentic-guard, an open-source static analyzer that detects confused-deputy and prompt-injection risks in LLM agent code by modeling the LLM as an adversarially-controlled edge in the taint graph.

Saurabh Yergattikar
Judge
Saurabh Yergattikar is a Lead Engineer / Member of Technical Staff-2 at eBay Inc. and a contributor to the open-source SAFE-MCP project (Linux Foundation / OpenSSF), as well as the architect and developer of the open-source ShieldMCP system.

Anchit Jhingan
Judge
Anchit Jhingan is a Senior Data Scientist with over 6 years of experience in big tech, specializing in machine learning, AI systems, and analytics. He currently works at Amazon Prime Video in the content localization domain. His work focuses on building intelligent systems to solve real-world business problems, particularly in media and entertainment.

Ashita Khetan
Judge
Ashita Khetan is a Principal Software Engineer at Microsoft with over 12 years of experience building large-scale enterprise and AI-powered productivity solutions used by hundreds of millions of users. She specializes in customer experience technologies, artificial intelligence, cloud platforms, and enterprise software, and contributes to the broader technology community through judging, peer review, and mentoring initiatives.

Spurthi Tallam
Judge
Spurthi is a senior machine learning engineer with seven years across research and production ML, currently building data, AI/ML systems at LePrix. She previously worked on LLM-powered conversational systems at Good Inside and on data and machine learning at Samsung Research, and holds an MS in Computer Science from UMass Amherst.

Surbhi Madan
Judge
Surbhi Madan is a Senior Software Engineer at Google working on the Google Maps AI rendering platforms and infrastructure (focused on AskMaps). She has worked at Google, based in NYC for 8+ years, and is passionate about creating scalable and sustainable platforms for feature teams. She also focuses on growing the next generation of tech talent by mentoring, teaching, and fostering a supportive and collaborative team environment. Surbhi is a graduate of Brown University and is heavily involved in Google's intern hiring program and has mentored several interns in the past.

Phani Harish Wajjala
Judge
Phani Harish Wajjala is a Principal Machine Learning Engineer at Roblox, where he leads the ML decisioning layer that classifies and moderates millions of user-generated 3D assets. His work spans content-moderation evaluation, multimodal/VLM pipelines, and model calibration at production scale, backed by five U.S. patents and a CVPR 2026 workshop publication in the area.

Pratham Patkar
Judge
Pratham Patkar is Director of Business Systems at Society for Science, where he leads enterprise data architecture, AI readiness strategy, and data governance across the organization's many initiatives — including STEM research competitions, science education outreach, and science news publishing. He has designed data governance frameworks addressing GDPR, CCPA, and COPPA compliance, and has published research on constituent-first data governance for mission-driven organizations. His work explores how nonprofits can adopt AI responsibly without compromising the trust and data protection obligations they hold toward vulnerable constituent populations.

Suneet Malhotra
Judge
Suneet Malhotra is an independent researcher in AI-augmented test automation with 20+ years in software quality engineering. His current work is on multi-agent SDLC orchestration, LLM-as-a-Judge evaluation, and cross-layer observability for agentic systems, with recent peer-reviewed submissions on these topics and open-source companion code at github.com/SuneetMalhotra.

FNU Tejinder
Judge
FNU Tejinder is a Senior Manager at Deloitte Consulting LLP with more than 22 years in supply chain planning systems and applied artificial intelligence for Fortune 500 manufacturers. His current work is on agentic AI systems that make planning decisions autonomously, and on the decision authority, auditability, and evaluation methods required for systems whose outputs are non-deterministic. He is an IEEE Senior Member and reviews for Elsevier journals including Engineering Applications of Artificial Intelligence, the NeurIPS Ethics Track, the ACM SIGKDD Workshop on Agentic AI Evaluation and Trustworthiness, and ACM Computing Reviews, where he is a Featured Reviewer.

Ankit Arya
Judge
Ankit is Head of AI at Inscope, where he builds AI-native financial reporting systems and works on the practical safety challenges of deploying LLMs in regulated, high-accuracy domains, from prompt injection defenses to evaluation design.

Akshay Iyer
Judge
CS and Entrepreneurship at Columbia University, IIT Bombay alum. Research experience in neuromorphic engineering and federated learning. Apart Research judge and internal collaborator.

Ashwin Pai
Judge
Ashwin Pai is an engineering leader with over a decade of experience building distributed systems and security products, most recently shipping AI governance and continuous control validation platforms to Fortune 500 customers at RelyanceAI. He was previously CTO of Interfold, a VC-backed fintech startup he built from zero to one.

Diego Gomez
Judge
Works on multimodal LLM safety at YouTube and is developing a mechanistic interpretability project on persona vectors at BlueDot Impact (https://bluedot.org).

Mateusz Jurewicz
Judge
Senior ML Engineer with over 10 years of industry experience and a PhD in Artificial Intelligence, currently leading a team of data scientists in the Agentic AI Department of a large financial institution. Interested in safe & universally beneficial AI through both research and application.

Siddhi Chaturvedi
Judge
Siddhi is a Software Engineer at Barclays, where she works on the bank's credit card services, building and maintaining backend systems that support customer accounts, payments, and secure financial transactions. She focuses on developing reliable APIs and scalable solutions that deliver secure, high-quality digital banking experiences.
Registered Local Sites
Register A Location
Beside the remote and virtual participation, our amazing organizers also host local hackathon locations where you can meet up in-person and connect with others in your area.
The in-person events for the Apart Sprints are run by passionate individuals just like you! We organize the schedule, speakers, and starter templates, and you can focus on engaging your local research, student, and engineering community.
Our Other Sprints
-
Research
AI Incident Response Sprint
This unique event brings together diverse perspectives to tackle crucial challenges in AI alignment, governance, and safety. Work alongside leading experts, develop innovative solutions, and help shape the future of responsible
Sign Up
-
Research
Secret Loyalties Hackathon
This unique event brings together diverse perspectives to tackle crucial challenges in AI alignment, governance, and safety. Work alongside leading experts, develop innovative solutions, and help shape the future of responsible
Sign Up

Sign up to stay updated on the
latest news, research, and events
Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Sign up to stay updated on the
latest news, research, and events
Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Sign up to stay updated on the
latest news, research, and events
Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Sign up to stay updated on the
latest news, research, and events
Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923
