MedOmitDetect: Ensuring Safety in Patient-Facing Medical LLM
Sonal Joshi
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Large language models (LLMs) show promise for patient-facing medical chatbots, but their deployment raises critical safety concerns. While existing evaluation frameworks focus on detecting factual inaccuracies (hallucinations), they often overlook a more insidious risk: omissions—missing critical information that patients cannot independently identify. We introduce MedOmitDetect, an automated evaluation system that assesses the completeness of medical AI responses by detecting clinically significant omissions and classifying them by potential patient harm severity (Mild, Moderate, Severe, Life-threatening). Built on the clinician-annotated MedExpert dataset & benchmark experiments, our system provides a two-stage pipeline for generating and evaluating medical LLM responses across 8 models, enabling systematic safety benchmarking before clinical deployment.
Reviews
No public critique yet.
Cite this project
@misc{joshi2025medomitdetect,
title = {{MedOmitDetect: Ensuring Safety in Patient-Facing Medical LLM}},
author = {Sonal Joshi},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/medomitdetect-ensuring-safety-in-patientfacing-medical-llm-k6q3}},
url = {https://apartresearch.com/sprints/projects/medomitdetect-ensuring-safety-in-patientfacing-medical-llm-k6q3}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …