Skip to content
Sprint projectAug 26, 2024

Unsolved AI Safety Concepts Explorer

Tewodros Mesfin · Team Theo

Submitted to AI capabilities and risks demo-jam. Sprint projects are early-stage work by participants, not Apart Research publications.

n interactive demonstration that showcases some unsolved fundamental AI safety concepts.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. I think explaining side is very weak: it feels very much just like “AI → bad” and doesn’t give intuitions for how things work.

  2. Nice presentation.I think this demo lacked a way to convey gear level understanding in the mispecifications and misalignment it wanted to. It was not clear in each failure mode presented what actually happened or why.

Cite this project

@misc{mesfin2024unsolved,
  title = {{Unsolved AI Safety Concepts Explorer}},
  author = {Tewodros Mesfin},
  year = {2024},
  month = aug,
  note = {Submitted to AI capabilities and risks demo-jam, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/unsolved-ai-safety-concepts-explorer}},
  url = {https://apartresearch.com/sprints/projects/unsolved-ai-safety-concepts-explorer}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026