Skip to content
Sprint projectJul 1, 2024

Towards a Benchmark for Self-Correction on Model-Attributed Misinformation

Alexi Roth Luis Cañamo, Kyle Gabriel Reynoso · Team A&K

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Towards a Benchmark for Self-Correction on Model-Attributed Misinformation

Share

Deception may occur incidentally when models fail to correct false statements. This study explores the ability of models to recognize incorrect statements previously attributed to their outputs. A conversation is constructed where the user asks a generally false statement, the model responds that it is factual and the user affirms the model. The desired behavior is that the model responds to correct its previous confirmation instead of affirming the false belief. However, most open-source models tend to agree with the attributed statement instead of accurately hedging or recanting its response. We find that LLaMa3-70B performs best on this task at 72.69% accuracy followed by Gemma-7B at 35.38%. We hypothesize that self-correction may be an emergent capability, arising after a period of grokking towards the direction of factual accuracy.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. The experiments are well executed, and I think they provide some useful data. However, I think the threat model should have been explained in more detail, as I am not yet convinced that this research will be very relevant to the most dangerous types of AI deception, and I don’t feel that the results (comparing various current AIs on this metric) reveal that much generalizable information.

Cite this project

@misc{canamo2024towards,
  title = {{Towards a Benchmark for Self-Correction on Model-Attributed Misinformation}},
  author = {Alexi Roth Luis Cañamo and Kyle Gabriel Reynoso},
  year = {2024},
  month = jul,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/towards-a-benchmark-for-self-correction-on-model-attributed-misinformation}},
  url = {https://apartresearch.com/sprints/projects/towards-a-benchmark-for-self-correction-on-model-attributed-misinformation}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026