Skip to content
Sprint projectApr 26, 2026Belgium

Automated Causal Graph Extraction and Value-of-Information Prioritization for AI Biorisk Modelling

Douw Marx

Submitted to AIxBio Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Automated Causal Graph Extraction and Value-of-Information Prioritization for AI Biorisk Modelling

Code (opens in new tab)
Share

Quantifying risk at scale requires defining and prioritizing hundreds of causal paths to harm. I present an automated pipeline that (i) extracts causal chains from a single source document with an LLM, (ii) collapses near-duplicate nodes using embeddings and paired merge-proposer / merge-validator LLMs, and (iii) elicits Beta and PERT priors per node. Nodes in the risk model are then ranked by betweenness centrality, Birnbaum importance, and Expected Value of Partial Perfect Information (EVPPI), after Monte Carlo sampling. The pipeline is demonstrated on biorisk and serves as a proof-of-concept that LLMs can prioritize risk and indicate where new evaluations would most change downstream decisions.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. - A limitation for discussion is that biorisk modeling can be significantly based on classified information or infohazards, which can reduce the overall applicability.

    - Expanding towards taking in multiple sources. In the real world, multiple sources would inform overall evaluation priorities. I was not quite sure from the methodology if that'd be possible.

    - Explain why evaluation prioritization is a current bottleneck in biorisk reduction.

    - A paragraph on how plausible your results were would have been helpful. Do the recommendations broadly align with overall evaluation priorities? You could discuss how human graders would judge the outputs of the automated scoring and the plausibility.

    - A control condition where you compare this method to asking an AI Agent directly to extract the information and differences in outcome would have been interesting.

  2. The problem area is neglected - we need more rigorous tooling for intervention prioritization in AIxBio, and this work addresses that gap. However, the intended audience and theory of change are unclear from the paper; there are some mentions of it in the limitations and future work section, but it should be clear for the beginning who this tool is directed at, how it should be used, what problem it solves and what the downstream implications are.

    “Granularity” is undefined (although the author notes this explicitly, which is appreciated). The paper would benefit from an operational definition and a sensitivity analysis. Using LLMs to elicit priors is risky, since LLMs are often overconfident and can produce numbers detached from reality, although I did not read the AutoEllicit paper and I can see how this is a necessary choice given the scope of the project. However, Gemini 2.5 Flash Lite is a weak choice for this, since it’s a step that is genuinely difficult for LLMs, which is why a frontier model (or several) should be used - I understand the budget constraints, but the paper would benefit from an acknowledgement here.

    Connected to that, I have a problem with how P(X | any parent active) is a single Beta but the actual conditional probabilities could vary substantially across the parents.. Either model parent-specific conditionals, keep the structure as a forest, or argue why the merged single-parameter approximation is acceptable. As written, the merge step invalidates the elicitation model. This is the most serious methodological issue I have with the paper.

    Standard betweenness sums over all s ≠ t, but in a risk DAG, only source-to-outcome paths are semantically meaningful. A source-to-outcome restricted betweenness (or flow betweenness) would be more appropriate.

    Notation is not properly introduced in the paper, which makes it hard to read and sometimes confusing, especially with some indexes colliding.

    Figure 2 is unreadable - the caption says "zoom for node labels", but those labels are unreadable at any zoom level. This is the only chance the reader has to see what the graph contains, especially with the pipeline being built in a way that does not allow easy replication.

    Read full reviewShow less

Cite this project

@misc{marx2026automated,
  title = {{Automated Causal Graph Extraction and Value-of-Information Prioritization for AI Biorisk Modelling}},
  author = {Douw Marx},
  year = {2026},
  month = apr,
  note = {Submitted to AIxBio Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/automated-causal-graph-extraction-and-valueofinformation-prioritization-for-ai-biorisk-modelling-l8fn}},
  url = {https://apartresearch.com/sprints/projects/automated-causal-graph-extraction-and-valueofinformation-prioritization-for-ai-biorisk-modelling-l8fn}
}
Browse all projects

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026