Skip to content
Sprint projectOct 27, 2024

Understanding Incentives To Build Uninterruptible Agentic AI Systems

Damin Curtis, M.A. International Affairs Norman Piotriowski, B.Sc. Data Science · Team Understanding Incentives To Build Uninterruptible Agentic AI Systems

Submitted to AI Policy Hackathon at Johns Hopkins University. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Understanding Incentives To Build Uninterruptible Agentic AI Systems

Presentation

Presentation: Understanding Incentives To Build Uninterruptible Agentic AI Systems

Share

This proposal addresses the development of agentic AI systems in the context of national security. While potentially beneficial, they pose significant risks if not aligned with human values. We argue that the increasing autonomy of AI necessitates robust analyses of interruptibility mechanisms, and whether there are scenarios where it is safer to omit them.

Key incentives for creating uninterruptible systems include perceived benefits from uninterrupted operations, low perceived risks to the controller, and fears of adversarial exploitation of shutdown options. Our proposal draws parallels to established systems like constitutions and mutual assured destruction strategies that maintain stability against changing values. In some instances this may be desirable, while in others it poses even greater risks than otherwise accepted.

To mitigate those risks, our proposal recommends implementing comprehensive monitoring to detect misalignment, establishing tiered access to interruption controls, and supporting research on managing adversarial AI threats. Overall, a proactive and multi-layered policy approach is essential to balance the transformative potential of agentic AI with necessary safety measures.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Very balanced and thoughtful considerations

  2. I thought this tackled an interesting issue that touches upon multiple realms of inquiry within the science of AI safety. Agentic systems are fairly novel but a shutdown feature would indupitably ensure against certain deceptive alignment issues and other problems of autonomous systems "going rouge"

Cite this project

@misc{curtis2024understanding,
  title = {{Understanding Incentives To Build Uninterruptible Agentic AI Systems}},
  author = {Damin Curtis and M.A. International Affairs Norman Piotriowski and B.Sc. Data Science},
  year = {2024},
  month = oct,
  note = {Submitted to AI Policy Hackathon at Johns Hopkins University, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/understanding-incentives-to-build-uninterruptible-agentic-ai-systems}},
  url = {https://apartresearch.com/sprints/projects/understanding-incentives-to-build-uninterruptible-agentic-ai-systems}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026