Amplified Wise Simulations for Safe Training and Deployment
Chris Leong
Submitted to Hackathon for Technical AI Safety Startups. Sprint projects are early-stage work by participants, not Apart Research publications.
Conflict of interest declaration: I advised Fazl on a funding request he was working on.
Re publishing: This PDF would require further modifications before publication.
I want to train (amplified) imitation agents of people who are wise to provide advice on navigating conflicting considerations when figuring out how to train and deploy AI safely.
Path to Impact: Train wise AI advisors -> organisations make better decisions about how to train and deploy AI -> safer AGI -> better outcomes for humanity
What is wisdom? Why focus on increasing wisdom? See image
Why use amplified imitation learning?
Attempting to train directly on wisdom suffers from the usual problems of the optimisation algorithm adversarially leveraging your blind spots, but worse because wisdom is an especially fuzzy concept.
Attempting to understand wisdom from a principled approach and build wise AI directly would require at least 50 years and iteration through multiple paradigms of research.
In contrast, if our objective is to imitation folk who are wise, we have a target that we can optimise hard on. Instead of using reinforcment learning to go beyond human level, we use amplification techniques like debate or iterated amplification.
How will these agents advise on decisions?
The humans will ultimately make the decisions. The agents don't have to directly tell the humans what to do, they simply have to inspire the humans to make better decisions. I expect that these agents will be most useful in helping humans figuring out how to navigate conflicting principles or frameworks.
Reviews
No public critique yet.
Cite this project
@misc{leong2024amplified,
title = {{Amplified Wise Simulations for Safe Training and Deployment}},
author = {Chris Leong},
year = {2024},
month = sep,
note = {Submitted to Hackathon for Technical AI Safety Startups, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/amplified-wise-simulations-for-safe-training-and-deployment}},
url = {https://apartresearch.com/sprints/projects/amplified-wise-simulations-for-safe-training-and-deployment}
}More from Hackathon for Technical AI Safety Startups
- 1st place by peer reviewView project: DarkForest - Defending the Authentic and Humane Web
DarkForest - Defending the Authentic and Humane Web
DarkForest
DarkForest is a pioneering Human Content Verification System (HCVS) designed to safeguard the authenticity of online spaces in the face of increasing AI-generated content. By leveraging graph-based reinforcement …
- View project: Jailbreaking general purpose robots
Jailbreaking general purpose robots
Luax Labs
We show that state of the art LLMs can be jailbroken by adversarial multimodal inputs, and that this can lead to dangerous scenarios if these LLMs are used as planners in robotics. We propose finetuning small multimodal …
- View project: AI Safety Collective - Crowdsourcing Solutions for Critical AI Safety Challenges
AI Safety Collective - Crowdsourcing Solutions for Critical AI Safety Challenges
AI Safety Collective
The AI Safety Collective is a global platform designed to enhance AI safety by crowdsourcing solutions to critical AI Safety challenges. As AI systems like large language models and multimodal systems become more …