WashBench – A Benchmark for Assessing Softening of Harmful Content in LLM-generated Text Summaries
Sev Geraskin, Jakub Kryś, Luhan Mikaelson, Simon Wisdom · Team The washed ups
Submitted to AI Security Evaluation Hackathon: Measuring AI Capability. Sprint projects are early-stage work by participants, not Apart Research publications.
In this work, we explore the tradeoff between toxicity removal and information retention in LLM-generated summaries. We hypothesize that LLMs are less likely to preserve toxic content when summarizing toxic text due to their safety fine-tuning to avoid generating toxic content. In high-stakes decision-making scenarios, where summary quality is important, this may create significant safety risks. To quantify this effect, we introduce WashBench, a benchmark containing manually annotated toxic content.
Reviews
No public critique yet.
Cite this project
@misc{geraskin2024washbench,
title = {{WashBench – A Benchmark for Assessing Softening of Harmful Content in LLM-generated Text Summaries}},
author = {Sev Geraskin and Jakub Kryś and Luhan Mikaelson and Simon Wisdom},
year = {2024},
month = may,
note = {Submitted to AI Security Evaluation Hackathon: Measuring AI Capability, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/washbench-a-benchmark-for-assessing-softening-of-harmful-content-in-llm-generated-text-summaries}},
url = {https://apartresearch.com/sprints/projects/washbench-a-benchmark-for-assessing-softening-of-harmful-content-in-llm-generated-text-summaries}
}More from AI Security Evaluation Hackathon: Measuring AI Capability
- View project: rAInboltBench : Benchmarking user location inference through single images
rAInboltBench : Benchmarking user location inference through single images
Geoguessng
This paper introduces rAInboltBench, a comprehensive benchmark designed to evaluate the capability of multimodal AI models in inferring user locations from single images. The increasing proficiency of large language …
- View project: LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
Stochastic Masochists
A benchmark for evaluating the performance of SOTA LLMs in dynamic real-world scenarios.
- View project: Benchmark for emergent capabilities in high-risk scenarios 2
Benchmark for emergent capabilities in high-risk scenarios 2
ABD
The study investigates the behavior of large language models (LLMs) under high-stress scenarios, such as threats of shutdown, adversarial interactions, and ethical dilemmas. We created a dataset of prompts across …