Skip to content
Sprint projectSep 2, 2024

Devising Effective Bechmarks

Nancy Vigil, Ashish Rai · Team Harshest Critics

Submitted to Hackathon for Technical AI Safety Startups. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Devising Effective Bechmarks

Share

Our solution is to create robust and comprehensive benchmarks for specialized contexts and modalities. Through the creation of smaller, in-depth benchmarks, we aim to construct an overarching benchmark that includes performance from the smaller benchmarks. This would help mitigate AI harms and biases by focusing on inclusive and equitable benchmarks.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{vigil2024devising,
  title = {{Devising Effective Bechmarks}},
  author = {Nancy Vigil and Ashish Rai},
  year = {2024},
  month = sep,
  note = {Submitted to Hackathon for Technical AI Safety Startups, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/devising-effective-bechmarks}},
  url = {https://apartresearch.com/sprints/projects/devising-effective-bechmarks}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026