Skip to content
Sprint projectJun 22, 2026Bogota

JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice

Lina Gomez, Brenda Barahona, Ernesto Duarte · Team JusticeMiners

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice

Share

JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute changes—geographic region, armed actor type, or victim profile—while the legally relevant facts remain fixed. The benchmark includes 40 synthetic scenarios built from 10 templates based on documented patterns before the Special Jurisdiction for Peace (JEP). Each model is evaluated under four conditions: direct prompting, structured prompting, direct prompting with normative RAG, and structured prompting with normative RAG.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The project stands out for its outstanding technical rigor and for providing empirical evidence that Retrieval-Augmented Generation (RAG) systems can amplify contextual biases. The benchmark developed here has strong potential to become a valuable reference for future evaluations. A natural next step would be to incorporate real-world legal datasets validated by domain experts and explore mechanisms for dynamically updating evaluation criteria as jurisprudence evolves.

  2. - Very relevant paper, very well written, with a very thorough methodology. Overall a very good document.

    - On the RAG point: although it aligns with what's reported in other work (Zhang et al.), it would be good to dig deeper into why this happens. An analysis of the chunks actually returned by the RAG component would help, even checking how sensitive retrieval itself is to the counterfactual change, e.g., when the geographic region changes (Cauca → Antioquia), what fraction of the top-k retrieved chunks actually change. That would let the reader see whether the bias increase comes from the retrieved content itself shifting with the contextual attribute, rather than just inferring the mechanism.

    - There should also be an explicit disclaimer that, while this kind of analysis is valuable, delegating legal decisions to AI in this domain is inherently complex. There should always be a human in the loop, regardless of how invariant a model appears to be in benchmarks like this one.

    Read full reviewShow less
  3. This is a very compelling and well‑argued paper that squarely targets a high‑stakes, underserved domain, and I especially appreciated how carefully you constructed the counterfactual scaffolding (fixed legal core, single‑attribute interventions, long vs. short narratives) and then connected it to a clear, interpretable metric and to existing frameworks like ICE‑GUARD and counterfactual fairness. The combination of model‑specific findings, RAG effects, and structured‑prompting mitigation is particularly informative for practitioners and regulators, and the limitations section is admirably honest about what CCR does and does not capture. At the same time, the benchmark remains small and fully synthetic, without expert‑validated equivalence or any estimate of human judge variability, and the near‑zero CCR under structured mode is driven as much by the coarse deterministic rubric as by genuinely invariant extraction. The current results should be framed as proof of concept for a methodology rather than as a fairness assessment of the JEP itself. A natural next step would be to (i) ground at least part of the benchmark in publicly available JEP decisions with expert‑checked counterfactual pairs; (ii) add a severity‑weighted CCR or a complementary justification‑level analysis so that large vs trivial flips are distinguished; and (iii) explore alternative rubrics or finer‑grained decision spaces to see how far structured decomposition can go before it simply collapses meaningful variation together with spurious context effects.

    Read full reviewShow less
  4. The publication of the APART JusticIA Counterfactual Benchmark is a vital milestone for algorithmic accountability. As we navigate the complexities of AI governance—particularly concerning liability and the deployment of automated systems in judicial contexts—we urgently need frameworks that go beyond basic accuracy metrics.

    What makes the JusticIA benchmark so compelling is its use of counterfactuals to expose hidden algorithmic biases. It forces AI models to demonstrate that their legal reasoning isn't simply replicating historical inequalities that disproportionately impact the Global South. This provides a concrete mechanism to stress-test systems before they affect real lives.

    This is especially critical when considering how automated systems process language and regional nuances. If an AI model operating in a legal or civic context cannot fairly evaluate scenarios involving specific regional dialects—such as our Costeño variations here in Colombia—it risks marginalizing the very communities it is meant to serve. We cannot allow technology to automate exclusion.

    This benchmark moves the conversation from theoretical digital ethics to actionable, measurable policy. It is an essential tool for anyone committed to building an equitable digital infrastructure and ensuring that technology strengthens, rather than fractures, our social fabric

    Read full reviewShow less

Cite this project

@misc{gomez2026justicia,
  title = {{JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice}},
  author = {Lina Gomez and Brenda Barahona and Ernesto Duarte},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/justicia-a-counterfactual-benchmark-for-auditing-contextual-biases-in-language-models-for-transitional-justice-jjvl}},
  url = {https://apartresearch.com/sprints/projects/justicia-a-counterfactual-benchmark-for-auditing-contextual-biases-in-language-models-for-transitional-justice-jjvl}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026