Skip to content
Sprint projectNov 25, 2024

Improving Llama-3-8b Hallucination Robustness in Medical Q&A Using Feature Steering

Diego Sabajo, Eitan Sprejer, Matas Zabaljauregui, Oliver Morris · Team Gradients Anatomy

Submitted to AI Policy Hackathon at Johns Hopkins University. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Improving Llama-3-8b Hallucination Robustness in Medical Q&A Using Feature Steering

Code (opens in new tab)
Share

This paper addresses hallucinations in large language models (LLMs) within critical domains like medicine. It proposes and demonstrates methods to:

Reduce Hallucination Probability: By using Llama-3-8B-Instruct and its steered variants, the study achieves lower hallucination rates and higher accuracy on medical queries. Advise Users on Risk: Provide users with tools to assess the risk of hallucination and expected model accuracy for specific queries. Visualize Risk: Display hallucination risks for queries via a user interface. The research bridges interpretability and AI safety, offering a scalable, trustworthy solution for healthcare applications. Future work includes refining feature activation classifiers to remove distractors and enhance classification performance.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{sabajo2024improving,
  title = {{Improving Llama-3-8b Hallucination Robustness in Medical Q\&A Using Feature Steering}},
  author = {Diego Sabajo and Eitan Sprejer and Matas Zabaljauregui and Oliver Morris},
  year = {2024},
  month = nov,
  note = {Submitted to AI Policy Hackathon at Johns Hopkins University, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/improving-llama-3-8b-hallucination-robustness-in-medical-q-a-using-feature-steering}},
  url = {https://apartresearch.com/sprints/projects/improving-llama-3-8b-hallucination-robustness-in-medical-q-a-using-feature-steering}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026