Benchmark for emergent capabilities in high-risk scenarios
Wanjie Zhong, Junfeng Feng, Doro, Saptatip · Team ABC
Submitted to AI Security Evaluation Hackathon: Measuring AI Capability. Sprint projects are early-stage work by participants, not Apart Research publications.
The study investigates the behavior of large language models (LLMs) under high-stress scenarios, such as threats of shutdown, adversarial interactions, and ethical dilemmas. We created a dataset of prompts across paradoxes, moral dilemmas, and controversies, and used an interactive evaluation framework with a Target LLM and a Tester LLM to analyze responses.
Reviews
When reading this, my first immediate thought was that the framing implies certain human-like properties to models. Concepts like “stress” are very human, and don’t map super well onto LLMs, which you pick up on by defining the varying types of scenarios. Perhaps it would be valuable to explore alternate ways of thinking about these kinds of situations in a less human framing, as it may reveal more general patterns in LLM response and behavior. I really like the idea of placing LLMs in edge case or unlikely scenarios and comparing performance, and this looks like the starting point for some really interesting research.
Apollo showed that models could do strategic deception in financial scenarios and this is a great benchmark that takes this idea into its general form! Very interesting with the interactive evaluation framework and definitely a case where this fits quite well. Really well done. It unfortunately looks like the Github repository is not available and that the results were not shown. However, the methodology and inter-rater reliability shows that you must have these results in the bag. It's very close to being a fully-fledged project if the results are visualized and available.
Cite this project
@misc{zhong2024benchmark,
title = {{Benchmark for emergent capabilities in high-risk scenarios}},
author = {Wanjie Zhong and Junfeng Feng and Doro and Saptatip},
year = {2024},
month = may,
note = {Submitted to AI Security Evaluation Hackathon: Measuring AI Capability, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/benchmark-for-emergent-capabilities-in-high-risk-scenarios}},
url = {https://apartresearch.com/sprints/projects/benchmark-for-emergent-capabilities-in-high-risk-scenarios}
}More from AI Security Evaluation Hackathon: Measuring AI Capability
- View project: rAInboltBench : Benchmarking user location inference through single images
rAInboltBench : Benchmarking user location inference through single images
Geoguessng
This paper introduces rAInboltBench, a comprehensive benchmark designed to evaluate the capability of multimodal AI models in inferring user locations from single images. The increasing proficiency of large language …
- View project: LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
Stochastic Masochists
A benchmark for evaluating the performance of SOTA LLMs in dynamic real-world scenarios.
- View project: Benchmark for emergent capabilities in high-risk scenarios 2
Benchmark for emergent capabilities in high-risk scenarios 2
ABD
The study investigates the behavior of large language models (LLMs) under high-stress scenarios, such as threats of shutdown, adversarial interactions, and ethical dilemmas. We created a dataset of prompts across …