Evaluating and inducing steganography in LLMs
Artem Karpov
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
This report demonstrates that large language models are capable of hiding simple 8 bit information in their output using associations from more powerful overseers (other LLMs or humans). Without direct steganography fine tuning, LLAMA 3 8B can guess a 8 bit hidden message in a plain text in most cases (69%), however a more capable model, GPT-3.5 was able to catch almost all of them (84%). More research is required to investigate how this ability might be induced or improved via RL training in similar and larger models.
Reviews
The work is the start of a fairly ambitious project to do PPO fine-tuning of LLMs for steganography (as the work acknowledges, this is notoriously hard). This is definitely an interesting direction, and I appreciate that someone has actually gone and taken a stab at the hard thing.
Cite this project
@misc{karpov2024evaluating,
title = {{Evaluating and inducing steganography in LLMs}},
author = {Artem Karpov},
year = {2024},
month = jun,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/evaluating-and-inducing-steganography-in-llms}},
url = {https://apartresearch.com/sprints/projects/evaluating-and-inducing-steganography-in-llms}
}More from Deception Detection Hackathon: Preventing AI deception
- 1st place by peer reviewView project: Sandbag Detection through Model Degradation
Sandbag Detection through Model Degradation
Team Truth Serum
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
- View project: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DeceptionRepE
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …