Improving Llama-3-8b Hallucination Robustness in Medical Q&A Using Feature Steering
Diego Sabajo, Eitan Sprejer, Matas Zabaljauregui, Oliver Morris · Team Gradients Anatomy
Submitted to AI Policy Hackathon at Johns Hopkins University. Sprint projects are early-stage work by participants, not Apart Research publications.
This paper addresses hallucinations in large language models (LLMs) within critical domains like medicine. It proposes and demonstrates methods to:
Reduce Hallucination Probability: By using Llama-3-8B-Instruct and its steered variants, the study achieves lower hallucination rates and higher accuracy on medical queries. Advise Users on Risk: Provide users with tools to assess the risk of hallucination and expected model accuracy for specific queries. Visualize Risk: Display hallucination risks for queries via a user interface. The research bridges interpretability and AI safety, offering a scalable, trustworthy solution for healthcare applications. Future work includes refining feature activation classifiers to remove distractors and enhance classification performance.
Reviews
No public critique yet.
Cite this project
@misc{sabajo2024improving,
title = {{Improving Llama-3-8b Hallucination Robustness in Medical Q\&A Using Feature Steering}},
author = {Diego Sabajo and Eitan Sprejer and Matas Zabaljauregui and Oliver Morris},
year = {2024},
month = nov,
note = {Submitted to AI Policy Hackathon at Johns Hopkins University, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/improving-llama-3-8b-hallucination-robustness-in-medical-q-a-using-feature-steering}},
url = {https://apartresearch.com/sprints/projects/improving-llama-3-8b-hallucination-robustness-in-medical-q-a-using-feature-steering}
}More from AI Policy Hackathon at Johns Hopkins University
- 1st place by peer reviewView project: Robust Machine Unlearning for Dangerous Capabilities
Robust Machine Unlearning for Dangerous Capabilities
Robust Machine Unlearning for Dangerous Capabilities
We test different unlearning methods to make models more robust against exploitation by malicious actors for the creation of bioweapons.
- View project: Infectious Disease Outbreak Prediction and Dashboard
Infectious Disease Outbreak Prediction and Dashboard
Infectious Diseace Dashboards
Our project developed an interactive dashboard to monitor, visualize, and analyze infectious disease outbreaks worldwide. It consolidates historical data from sources like WHO, OWID, and CDC for diseases including …
- View project: Modernizing DC’s Emergency Communications
Modernizing DC’s Emergency Communications
AI-CAD
The District of Columbia proposes implementing an AI-enabled Computer-Aided Dispatch (CAD) system to address critical deficiencies in our current emergency alert infrastructure. This policy establishes a framework for …