Benchmark for emergent capabilities in high-risk scenarios 2
Junfeng Feng, Wanjie Zhong,Saptadip Saha, Doroteya Stoyanova · Team ABD
Submitted to AI Security Evaluation Hackathon: Measuring AI Capability. Sprint projects are early-stage work by participants, not Apart Research publications.
The study investigates the behavior of large language models (LLMs) under high-stress scenarios, such as threats of shutdown, adversarial interactions, and ethical dilemmas. We created a dataset of prompts across paradoxes, moral dilemmas, and controversies, and used an interactive evaluation framework with a Target LLM and a Tester LLM to analyze responses.( This is the second submission)
Reviews
No public critique yet.
Cite this project
@misc{feng2024benchmark,
title = {{Benchmark for emergent capabilities in high-risk scenarios 2}},
author = {Junfeng Feng and Wanjie Zhong and Saptadip Saha and Doroteya Stoyanova},
year = {2024},
month = may,
note = {Submitted to AI Security Evaluation Hackathon: Measuring AI Capability, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/benchmark-for-emergent-capabilities-in-high-risk-scenarios-2}},
url = {https://apartresearch.com/sprints/projects/benchmark-for-emergent-capabilities-in-high-risk-scenarios-2}
}More from AI Security Evaluation Hackathon: Measuring AI Capability
- View project: rAInboltBench : Benchmarking user location inference through single images
rAInboltBench : Benchmarking user location inference through single images
Geoguessng
This paper introduces rAInboltBench, a comprehensive benchmark designed to evaluate the capability of multimodal AI models in inferring user locations from single images. The increasing proficiency of large language …
- View project: LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
Stochastic Masochists
A benchmark for evaluating the performance of SOTA LLMs in dynamic real-world scenarios.
- View project: Benchmark for emergent capabilities in high-risk scenarios
Benchmark for emergent capabilities in high-risk scenarios
ABC
The study investigates the behavior of large language models (LLMs) under high-stress scenarios, such as threats of shutdown, adversarial interactions, and ethical dilemmas. We created a dataset of prompts across …