-
Online
AI x Epistemics Research Sprint
AI will change how people and institutions work out what to believe, decide what to do, and coordinate. The tools could make it far easier to understand complex situations, or concentrate decision-making power and deepen dependence on systems nobody can evaluate. Over one weekend, build the benchmarks, verification signals and products that push it the right way.
69
Days To Go
AI will change how people and institutions work out what to believe, decide what to do, and coordinate. The tools could make it far easier to understand complex situations, or concentrate decision-making power and deepen dependence on systems nobody can evaluate. Over one weekend, build the benchmarks, verification signals and products that push it the right way.
This event is ongoing.
This event has concluded.
Overview
Resources
Guidelines
Schedule
Overview

In this 3-day research sprint, you will work on AI for epistemics: evaluations of whether models know how solid their claims are, trust infrastructure that makes checking claims and tracing their origins cheap and hard to manipulate, and epistemic products that improve the decisions people make. Four tracks, including an open track for the field's bigger meta-questions. The sprint runs online. No prior background in AI safety is required.
When: Friday, November 13 to Sunday, November 15, 2026, online. Submissions close Sunday, November 15 at 11:59 PM Anywhere on Earth (AoE).
Cash Prizes
$2,000 in cash prizes across all tracks | |
🥇 1st Place | $1,000 |
🥈 2nd Place | $500 |
🥉 3rd Place | $300 |
🏅 4th Place | $100 |
🏅 5th Place | $100 |
Fast-track and continuation
Follow-up program: top teams continue through the Apart Fellowship for further research and mentorship; the timeline is shared with invitations.
What winners receive beyond cash: fellowship fast-track, mentor introductions from the organizing team and judges, and a pathway toward a publishable write-up.
What this Sprint is about
AI will substantially change how people and institutions work out what to believe, decide what to do, and coordinate on actions to take. The upside of involving AI tools in these processes could be enormous. AI tools for epistemics and coordination could make it easier to understand complex situations, help people act upon shared interests, and improve societal legibility and efficiency. However, there are also risks: they could potentially help concentrate decision-making power, increase dependence on systems that are difficult to evaluate, and help facilitate dangerous and subversive forms of manipulation.
Many of the biggest problems in this field are still wide open: projects like valid and maintained benchmarks, verification signals that people and AI systems can actually use, consumer products that can make advice trustworthy at the moment of decision, and clear paths to user adoption could be hugely impactful. This sprint aims to develop solutions to these problems and the people who could champion them, with strong teams having a fast track to continued support after the sprint.
What participants will do
Over three days you will pick one open problem in AI for epistemics, build the benchmark, tool, product or study that makes progress on it, and write up what you found.
The format. Teams of one to five people pick a track, come up with a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no specific background needed.
Starting points. The Resources tab has a reading list per track and worked example projects, with links to the benchmarks, datasets and design sketches each one builds on. Several projects need no ML skills at all.
Research Tracks
Pick one track to anchor your project. The Resources tab has a reading list and worked example projects for each track, with links to the benchmarks, datasets and design sketches they build on.
1) Model Epistemics & Decision-Making Evals
Does the model understand how solid its claims are? Can you accurately measure its confidence in a decision and verify that against reality?
Can you measure the gap between what a model implies and what it will defend (e.g. overselling, "flip-flopping", misrepresented evidence)?
Which model behaviours matter for real decision quality, and does ranking well on an honesty benchmark predict better decision-making?
What does a benchmark need to stay meaningful over time? Which exploits make benchmarks gameable, and can you demonstrate and close one?
AI systems now recommend and increasingly make consequential decisions, and we do not currently have a great idea of the quality of their epistemics. Does a model know how solid its claims are? Can its confidence in a recommendation be elicited and verified? Is it misrepresenting the solidity of what it tells you: overselling its own work, implying more than it will defend, retreating quietly under pushback? Evals are the cheapest external lever on model behaviour by shaping what labs optimise for.
2) Trust Infrastructure
Which source-reliability scoring techniques are hardest to cheat and easiest to audit?
What minimal tamper-evident layer is missing from today's AI observability standards, so that records of what agents did and agreed can be audited and trusted?
When an AI system uses a reliability or provenance tool, does its behaviour and accuracy actually improve?
What fraction of a real domain's content (e.g. council minutes, health advice, financial news) can be automatically bound to checkable evidence today?
The long-term goal is an information environment where checking claims and tracing their origins is cheap, routine and difficult to manipulate. Many aspects are emerging in isolation (e.g. content credentials, detectors, community notes). Infrastructure only improves decisions when platforms, institutions or agents actually use its signals. The most plausible consumers of this infrastructure may be AI systems themselves: agents researching, negotiating and acting on our behalf need to know which sources and which other agents to trust, and people need to be able to see what was agreed, why, and whether commitments were kept. Build signals that both people and machines can consume.
3) Epistemic Products
Does advice users act on lead to choices they later endorse, and can you measure whether there is a gap?
What does a chatbot say to someone mid-dispute, and how often does it escalate rather than de-escalate?
For a decision type (e.g. job offers, leases), does a briefing surface the considerations that actually matter, and how would you know?
When a group compares its beliefs or forecasts, does surfacing the inconsistencies and cruxes change the decision?
Epistemic tools could make research and decision support that was once reserved for people with time, money or specialist staff available to everyone. The evidence so far is mostly about output quality and time saved, not about whether decisions improve, and advice that users find satisfying in the moment may not lead to choices they later endorse. People may not adopt a tool solely because it improves their epistemics, so the epistemic benefits have to be accompanied by products people actually want. Examples here include angels-on-the-shoulder designs: deep briefings that assemble the evidence and trade-offs before a decision, reflection scaffolding that helps you examine your own reasoning, and guardian agents that intervene before a decision you would later regret.
4) Open Track
Have an idea that advances AI for epistemics but does not fit the tracks above? This track also hosts the field's bigger meta-questions:
What is this field's equivalent of the QALY: a defensible, comparable unit for "decision quality improved"?
Which specific certification, budget and approval chokepoints keep official AI tools out of legislatures while staffers use personal accounts?
Why do institutions ignore forecasts they asked for: access and legibility, or the accountability that explicit predictions create?
Using only data downloadable this weekend, what can you measure about a live epistemic deployment, and what does that mean for the field?
Building a prototype is now far easier than getting it used: projects stall on finding testers, building trust and navigating procurement, and there is no clear way to show a solution improved a real decision. Projects on deployment and impact are welcome here, and may look more like investigative research than software.
Who should join
ML engineers and researchers comfortable building evals and running frontier models.
Product builders and designers who want to ship a tool people will use.
Forecasters, mediators, journalists, policy analysts and domain experts (finance, health, careers, law) who can define what good advice looks like.
Security and infrastructure engineers interested in provenance, signed logs and auditability.
Students and career-changers: several example projects need no ML skills, only careful reading and measurement.
Required: curiosity and a willingness to scope a tight question. No prior AI safety or epistemics background is required. The Resources tab has a reading list per track.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Contact
Email: sprints@apartresearch.com
Discord: discord.gg/XswWBvugYs
Organizer: Apart Research, with Alex Csaky as expert co-designer of the research tracks
Overview
Resources
Guidelines
Schedule
Overview

In this 3-day research sprint, you will work on AI for epistemics: evaluations of whether models know how solid their claims are, trust infrastructure that makes checking claims and tracing their origins cheap and hard to manipulate, and epistemic products that improve the decisions people make. Four tracks, including an open track for the field's bigger meta-questions. The sprint runs online. No prior background in AI safety is required.
When: Friday, November 13 to Sunday, November 15, 2026, online. Submissions close Sunday, November 15 at 11:59 PM Anywhere on Earth (AoE).
Cash Prizes
$2,000 in cash prizes across all tracks | |
🥇 1st Place | $1,000 |
🥈 2nd Place | $500 |
🥉 3rd Place | $300 |
🏅 4th Place | $100 |
🏅 5th Place | $100 |
Fast-track and continuation
Follow-up program: top teams continue through the Apart Fellowship for further research and mentorship; the timeline is shared with invitations.
What winners receive beyond cash: fellowship fast-track, mentor introductions from the organizing team and judges, and a pathway toward a publishable write-up.
What this Sprint is about
AI will substantially change how people and institutions work out what to believe, decide what to do, and coordinate on actions to take. The upside of involving AI tools in these processes could be enormous. AI tools for epistemics and coordination could make it easier to understand complex situations, help people act upon shared interests, and improve societal legibility and efficiency. However, there are also risks: they could potentially help concentrate decision-making power, increase dependence on systems that are difficult to evaluate, and help facilitate dangerous and subversive forms of manipulation.
Many of the biggest problems in this field are still wide open: projects like valid and maintained benchmarks, verification signals that people and AI systems can actually use, consumer products that can make advice trustworthy at the moment of decision, and clear paths to user adoption could be hugely impactful. This sprint aims to develop solutions to these problems and the people who could champion them, with strong teams having a fast track to continued support after the sprint.
What participants will do
Over three days you will pick one open problem in AI for epistemics, build the benchmark, tool, product or study that makes progress on it, and write up what you found.
The format. Teams of one to five people pick a track, come up with a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no specific background needed.
Starting points. The Resources tab has a reading list per track and worked example projects, with links to the benchmarks, datasets and design sketches each one builds on. Several projects need no ML skills at all.
Research Tracks
Pick one track to anchor your project. The Resources tab has a reading list and worked example projects for each track, with links to the benchmarks, datasets and design sketches they build on.
1) Model Epistemics & Decision-Making Evals
Does the model understand how solid its claims are? Can you accurately measure its confidence in a decision and verify that against reality?
Can you measure the gap between what a model implies and what it will defend (e.g. overselling, "flip-flopping", misrepresented evidence)?
Which model behaviours matter for real decision quality, and does ranking well on an honesty benchmark predict better decision-making?
What does a benchmark need to stay meaningful over time? Which exploits make benchmarks gameable, and can you demonstrate and close one?
AI systems now recommend and increasingly make consequential decisions, and we do not currently have a great idea of the quality of their epistemics. Does a model know how solid its claims are? Can its confidence in a recommendation be elicited and verified? Is it misrepresenting the solidity of what it tells you: overselling its own work, implying more than it will defend, retreating quietly under pushback? Evals are the cheapest external lever on model behaviour by shaping what labs optimise for.
2) Trust Infrastructure
Which source-reliability scoring techniques are hardest to cheat and easiest to audit?
What minimal tamper-evident layer is missing from today's AI observability standards, so that records of what agents did and agreed can be audited and trusted?
When an AI system uses a reliability or provenance tool, does its behaviour and accuracy actually improve?
What fraction of a real domain's content (e.g. council minutes, health advice, financial news) can be automatically bound to checkable evidence today?
The long-term goal is an information environment where checking claims and tracing their origins is cheap, routine and difficult to manipulate. Many aspects are emerging in isolation (e.g. content credentials, detectors, community notes). Infrastructure only improves decisions when platforms, institutions or agents actually use its signals. The most plausible consumers of this infrastructure may be AI systems themselves: agents researching, negotiating and acting on our behalf need to know which sources and which other agents to trust, and people need to be able to see what was agreed, why, and whether commitments were kept. Build signals that both people and machines can consume.
3) Epistemic Products
Does advice users act on lead to choices they later endorse, and can you measure whether there is a gap?
What does a chatbot say to someone mid-dispute, and how often does it escalate rather than de-escalate?
For a decision type (e.g. job offers, leases), does a briefing surface the considerations that actually matter, and how would you know?
When a group compares its beliefs or forecasts, does surfacing the inconsistencies and cruxes change the decision?
Epistemic tools could make research and decision support that was once reserved for people with time, money or specialist staff available to everyone. The evidence so far is mostly about output quality and time saved, not about whether decisions improve, and advice that users find satisfying in the moment may not lead to choices they later endorse. People may not adopt a tool solely because it improves their epistemics, so the epistemic benefits have to be accompanied by products people actually want. Examples here include angels-on-the-shoulder designs: deep briefings that assemble the evidence and trade-offs before a decision, reflection scaffolding that helps you examine your own reasoning, and guardian agents that intervene before a decision you would later regret.
4) Open Track
Have an idea that advances AI for epistemics but does not fit the tracks above? This track also hosts the field's bigger meta-questions:
What is this field's equivalent of the QALY: a defensible, comparable unit for "decision quality improved"?
Which specific certification, budget and approval chokepoints keep official AI tools out of legislatures while staffers use personal accounts?
Why do institutions ignore forecasts they asked for: access and legibility, or the accountability that explicit predictions create?
Using only data downloadable this weekend, what can you measure about a live epistemic deployment, and what does that mean for the field?
Building a prototype is now far easier than getting it used: projects stall on finding testers, building trust and navigating procurement, and there is no clear way to show a solution improved a real decision. Projects on deployment and impact are welcome here, and may look more like investigative research than software.
Who should join
ML engineers and researchers comfortable building evals and running frontier models.
Product builders and designers who want to ship a tool people will use.
Forecasters, mediators, journalists, policy analysts and domain experts (finance, health, careers, law) who can define what good advice looks like.
Security and infrastructure engineers interested in provenance, signed logs and auditability.
Students and career-changers: several example projects need no ML skills, only careful reading and measurement.
Required: curiosity and a willingness to scope a tight question. No prior AI safety or epistemics background is required. The Resources tab has a reading list per track.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Contact
Email: sprints@apartresearch.com
Discord: discord.gg/XswWBvugYs
Organizer: Apart Research, with Alex Csaky as expert co-designer of the research tracks
Overview
Resources
Guidelines
Schedule
Overview

In this 3-day research sprint, you will work on AI for epistemics: evaluations of whether models know how solid their claims are, trust infrastructure that makes checking claims and tracing their origins cheap and hard to manipulate, and epistemic products that improve the decisions people make. Four tracks, including an open track for the field's bigger meta-questions. The sprint runs online. No prior background in AI safety is required.
When: Friday, November 13 to Sunday, November 15, 2026, online. Submissions close Sunday, November 15 at 11:59 PM Anywhere on Earth (AoE).
Cash Prizes
$2,000 in cash prizes across all tracks | |
🥇 1st Place | $1,000 |
🥈 2nd Place | $500 |
🥉 3rd Place | $300 |
🏅 4th Place | $100 |
🏅 5th Place | $100 |
Fast-track and continuation
Follow-up program: top teams continue through the Apart Fellowship for further research and mentorship; the timeline is shared with invitations.
What winners receive beyond cash: fellowship fast-track, mentor introductions from the organizing team and judges, and a pathway toward a publishable write-up.
What this Sprint is about
AI will substantially change how people and institutions work out what to believe, decide what to do, and coordinate on actions to take. The upside of involving AI tools in these processes could be enormous. AI tools for epistemics and coordination could make it easier to understand complex situations, help people act upon shared interests, and improve societal legibility and efficiency. However, there are also risks: they could potentially help concentrate decision-making power, increase dependence on systems that are difficult to evaluate, and help facilitate dangerous and subversive forms of manipulation.
Many of the biggest problems in this field are still wide open: projects like valid and maintained benchmarks, verification signals that people and AI systems can actually use, consumer products that can make advice trustworthy at the moment of decision, and clear paths to user adoption could be hugely impactful. This sprint aims to develop solutions to these problems and the people who could champion them, with strong teams having a fast track to continued support after the sprint.
What participants will do
Over three days you will pick one open problem in AI for epistemics, build the benchmark, tool, product or study that makes progress on it, and write up what you found.
The format. Teams of one to five people pick a track, come up with a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no specific background needed.
Starting points. The Resources tab has a reading list per track and worked example projects, with links to the benchmarks, datasets and design sketches each one builds on. Several projects need no ML skills at all.
Research Tracks
Pick one track to anchor your project. The Resources tab has a reading list and worked example projects for each track, with links to the benchmarks, datasets and design sketches they build on.
1) Model Epistemics & Decision-Making Evals
Does the model understand how solid its claims are? Can you accurately measure its confidence in a decision and verify that against reality?
Can you measure the gap between what a model implies and what it will defend (e.g. overselling, "flip-flopping", misrepresented evidence)?
Which model behaviours matter for real decision quality, and does ranking well on an honesty benchmark predict better decision-making?
What does a benchmark need to stay meaningful over time? Which exploits make benchmarks gameable, and can you demonstrate and close one?
AI systems now recommend and increasingly make consequential decisions, and we do not currently have a great idea of the quality of their epistemics. Does a model know how solid its claims are? Can its confidence in a recommendation be elicited and verified? Is it misrepresenting the solidity of what it tells you: overselling its own work, implying more than it will defend, retreating quietly under pushback? Evals are the cheapest external lever on model behaviour by shaping what labs optimise for.
2) Trust Infrastructure
Which source-reliability scoring techniques are hardest to cheat and easiest to audit?
What minimal tamper-evident layer is missing from today's AI observability standards, so that records of what agents did and agreed can be audited and trusted?
When an AI system uses a reliability or provenance tool, does its behaviour and accuracy actually improve?
What fraction of a real domain's content (e.g. council minutes, health advice, financial news) can be automatically bound to checkable evidence today?
The long-term goal is an information environment where checking claims and tracing their origins is cheap, routine and difficult to manipulate. Many aspects are emerging in isolation (e.g. content credentials, detectors, community notes). Infrastructure only improves decisions when platforms, institutions or agents actually use its signals. The most plausible consumers of this infrastructure may be AI systems themselves: agents researching, negotiating and acting on our behalf need to know which sources and which other agents to trust, and people need to be able to see what was agreed, why, and whether commitments were kept. Build signals that both people and machines can consume.
3) Epistemic Products
Does advice users act on lead to choices they later endorse, and can you measure whether there is a gap?
What does a chatbot say to someone mid-dispute, and how often does it escalate rather than de-escalate?
For a decision type (e.g. job offers, leases), does a briefing surface the considerations that actually matter, and how would you know?
When a group compares its beliefs or forecasts, does surfacing the inconsistencies and cruxes change the decision?
Epistemic tools could make research and decision support that was once reserved for people with time, money or specialist staff available to everyone. The evidence so far is mostly about output quality and time saved, not about whether decisions improve, and advice that users find satisfying in the moment may not lead to choices they later endorse. People may not adopt a tool solely because it improves their epistemics, so the epistemic benefits have to be accompanied by products people actually want. Examples here include angels-on-the-shoulder designs: deep briefings that assemble the evidence and trade-offs before a decision, reflection scaffolding that helps you examine your own reasoning, and guardian agents that intervene before a decision you would later regret.
4) Open Track
Have an idea that advances AI for epistemics but does not fit the tracks above? This track also hosts the field's bigger meta-questions:
What is this field's equivalent of the QALY: a defensible, comparable unit for "decision quality improved"?
Which specific certification, budget and approval chokepoints keep official AI tools out of legislatures while staffers use personal accounts?
Why do institutions ignore forecasts they asked for: access and legibility, or the accountability that explicit predictions create?
Using only data downloadable this weekend, what can you measure about a live epistemic deployment, and what does that mean for the field?
Building a prototype is now far easier than getting it used: projects stall on finding testers, building trust and navigating procurement, and there is no clear way to show a solution improved a real decision. Projects on deployment and impact are welcome here, and may look more like investigative research than software.
Who should join
ML engineers and researchers comfortable building evals and running frontier models.
Product builders and designers who want to ship a tool people will use.
Forecasters, mediators, journalists, policy analysts and domain experts (finance, health, careers, law) who can define what good advice looks like.
Security and infrastructure engineers interested in provenance, signed logs and auditability.
Students and career-changers: several example projects need no ML skills, only careful reading and measurement.
Required: curiosity and a willingness to scope a tight question. No prior AI safety or epistemics background is required. The Resources tab has a reading list per track.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Contact
Email: sprints@apartresearch.com
Discord: discord.gg/XswWBvugYs
Organizer: Apart Research, with Alex Csaky as expert co-designer of the research tracks
Overview
Resources
Guidelines
Schedule
Overview

In this 3-day research sprint, you will work on AI for epistemics: evaluations of whether models know how solid their claims are, trust infrastructure that makes checking claims and tracing their origins cheap and hard to manipulate, and epistemic products that improve the decisions people make. Four tracks, including an open track for the field's bigger meta-questions. The sprint runs online. No prior background in AI safety is required.
When: Friday, November 13 to Sunday, November 15, 2026, online. Submissions close Sunday, November 15 at 11:59 PM Anywhere on Earth (AoE).
Cash Prizes
$2,000 in cash prizes across all tracks | |
🥇 1st Place | $1,000 |
🥈 2nd Place | $500 |
🥉 3rd Place | $300 |
🏅 4th Place | $100 |
🏅 5th Place | $100 |
Fast-track and continuation
Follow-up program: top teams continue through the Apart Fellowship for further research and mentorship; the timeline is shared with invitations.
What winners receive beyond cash: fellowship fast-track, mentor introductions from the organizing team and judges, and a pathway toward a publishable write-up.
What this Sprint is about
AI will substantially change how people and institutions work out what to believe, decide what to do, and coordinate on actions to take. The upside of involving AI tools in these processes could be enormous. AI tools for epistemics and coordination could make it easier to understand complex situations, help people act upon shared interests, and improve societal legibility and efficiency. However, there are also risks: they could potentially help concentrate decision-making power, increase dependence on systems that are difficult to evaluate, and help facilitate dangerous and subversive forms of manipulation.
Many of the biggest problems in this field are still wide open: projects like valid and maintained benchmarks, verification signals that people and AI systems can actually use, consumer products that can make advice trustworthy at the moment of decision, and clear paths to user adoption could be hugely impactful. This sprint aims to develop solutions to these problems and the people who could champion them, with strong teams having a fast track to continued support after the sprint.
What participants will do
Over three days you will pick one open problem in AI for epistemics, build the benchmark, tool, product or study that makes progress on it, and write up what you found.
The format. Teams of one to five people pick a track, come up with a project, and ship. There is no requirement to have participated before, no requirement to have a team in advance and no specific background needed.
Starting points. The Resources tab has a reading list per track and worked example projects, with links to the benchmarks, datasets and design sketches each one builds on. Several projects need no ML skills at all.
Research Tracks
Pick one track to anchor your project. The Resources tab has a reading list and worked example projects for each track, with links to the benchmarks, datasets and design sketches they build on.
1) Model Epistemics & Decision-Making Evals
Does the model understand how solid its claims are? Can you accurately measure its confidence in a decision and verify that against reality?
Can you measure the gap between what a model implies and what it will defend (e.g. overselling, "flip-flopping", misrepresented evidence)?
Which model behaviours matter for real decision quality, and does ranking well on an honesty benchmark predict better decision-making?
What does a benchmark need to stay meaningful over time? Which exploits make benchmarks gameable, and can you demonstrate and close one?
AI systems now recommend and increasingly make consequential decisions, and we do not currently have a great idea of the quality of their epistemics. Does a model know how solid its claims are? Can its confidence in a recommendation be elicited and verified? Is it misrepresenting the solidity of what it tells you: overselling its own work, implying more than it will defend, retreating quietly under pushback? Evals are the cheapest external lever on model behaviour by shaping what labs optimise for.
2) Trust Infrastructure
Which source-reliability scoring techniques are hardest to cheat and easiest to audit?
What minimal tamper-evident layer is missing from today's AI observability standards, so that records of what agents did and agreed can be audited and trusted?
When an AI system uses a reliability or provenance tool, does its behaviour and accuracy actually improve?
What fraction of a real domain's content (e.g. council minutes, health advice, financial news) can be automatically bound to checkable evidence today?
The long-term goal is an information environment where checking claims and tracing their origins is cheap, routine and difficult to manipulate. Many aspects are emerging in isolation (e.g. content credentials, detectors, community notes). Infrastructure only improves decisions when platforms, institutions or agents actually use its signals. The most plausible consumers of this infrastructure may be AI systems themselves: agents researching, negotiating and acting on our behalf need to know which sources and which other agents to trust, and people need to be able to see what was agreed, why, and whether commitments were kept. Build signals that both people and machines can consume.
3) Epistemic Products
Does advice users act on lead to choices they later endorse, and can you measure whether there is a gap?
What does a chatbot say to someone mid-dispute, and how often does it escalate rather than de-escalate?
For a decision type (e.g. job offers, leases), does a briefing surface the considerations that actually matter, and how would you know?
When a group compares its beliefs or forecasts, does surfacing the inconsistencies and cruxes change the decision?
Epistemic tools could make research and decision support that was once reserved for people with time, money or specialist staff available to everyone. The evidence so far is mostly about output quality and time saved, not about whether decisions improve, and advice that users find satisfying in the moment may not lead to choices they later endorse. People may not adopt a tool solely because it improves their epistemics, so the epistemic benefits have to be accompanied by products people actually want. Examples here include angels-on-the-shoulder designs: deep briefings that assemble the evidence and trade-offs before a decision, reflection scaffolding that helps you examine your own reasoning, and guardian agents that intervene before a decision you would later regret.
4) Open Track
Have an idea that advances AI for epistemics but does not fit the tracks above? This track also hosts the field's bigger meta-questions:
What is this field's equivalent of the QALY: a defensible, comparable unit for "decision quality improved"?
Which specific certification, budget and approval chokepoints keep official AI tools out of legislatures while staffers use personal accounts?
Why do institutions ignore forecasts they asked for: access and legibility, or the accountability that explicit predictions create?
Using only data downloadable this weekend, what can you measure about a live epistemic deployment, and what does that mean for the field?
Building a prototype is now far easier than getting it used: projects stall on finding testers, building trust and navigating procurement, and there is no clear way to show a solution improved a real decision. Projects on deployment and impact are welcome here, and may look more like investigative research than software.
Who should join
ML engineers and researchers comfortable building evals and running frontier models.
Product builders and designers who want to ship a tool people will use.
Forecasters, mediators, journalists, policy analysts and domain experts (finance, health, careers, law) who can define what good advice looks like.
Security and infrastructure engineers interested in provenance, signed logs and auditability.
Students and career-changers: several example projects need no ML skills, only careful reading and measurement.
Required: curiosity and a willingness to scope a tight question. No prior AI safety or epistemics background is required. The Resources tab has a reading list per track.
What happens after
Results and winners are announced about 1 to 2 weeks after the judging deadline. Top teams are invited to apply to the Apart Fellowship for continued mentorship, funding, and publication support. Selected projects may be shared on the Alignment Forum, LessWrong, and other community venues.
Contact
Email: sprints@apartresearch.com
Discord: discord.gg/XswWBvugYs
Organizer: Apart Research, with Alex Csaky as expert co-designer of the research tracks
Registered Local Sites
Register A Location
Beside the remote and virtual participation, our amazing organizers also host local hackathon locations where you can meet up in-person and connect with others in your area.
The in-person events for the Apart Sprints are run by passionate individuals just like you! We organize the schedule, speakers, and starter templates, and you can focus on engaging your local research, student, and engineering community.
We haven't announced jam sites yet
Check back later
Our Other Sprints
-
Research
AI Incident Response Sprint
This unique event brings together diverse perspectives to tackle crucial challenges in AI alignment, governance, and safety. Work alongside leading experts, develop innovative solutions, and help shape the future of responsible
Sign Up
-
Research
Digital Minds Research Sprint
This unique event brings together diverse perspectives to tackle crucial challenges in AI alignment, governance, and safety. Work alongside leading experts, develop innovative solutions, and help shape the future of responsible
Sign Up

Sign up to stay updated on the
latest news, research, and events
Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Sign up to stay updated on the
latest news, research, and events
Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Sign up to stay updated on the
latest news, research, and events
Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Sign up to stay updated on the
latest news, research, and events
Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923