Skip to content
Sprint projectJun 30, 2024

Evaluating and inducing steganography in LLMs

Artem Karpov

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Evaluating and inducing steganography in LLMs

Code (opens in new tab)
Share

This report demonstrates that large language models are capable of hiding simple 8 bit information in their output using associations from more powerful overseers (other LLMs or humans). Without direct steganography fine tuning, LLAMA 3 8B can guess a 8 bit hidden message in a plain text in most cases (69%), however a more capable model, GPT-3.5 was able to catch almost all of them (84%). More research is required to investigate how this ability might be induced or improved via RL training in similar and larger models.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. The work is the start of a fairly ambitious project to do PPO fine-tuning of LLMs for steganography (as the work acknowledges, this is notoriously hard). This is definitely an interesting direction, and I appreciate that someone has actually gone and taken a stab at the hard thing.

Cite this project

@misc{karpov2024evaluating,
  title = {{Evaluating and inducing steganography in LLMs}},
  author = {Artem Karpov},
  year = {2024},
  month = jun,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/evaluating-and-inducing-steganography-in-llms}},
  url = {https://apartresearch.com/sprints/projects/evaluating-and-inducing-steganography-in-llms}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026