Skip to content
Sprint projectJun 22, 2026India

GSM-PathEval: A Global South Robustness Benchmark for Telepathology Vision-Language Models

Anshuman Awasthi · Team AnshumanAI

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: GSM-PathEval: A Global South Robustness Benchmark for Telepathology Vision-Language Models

Code (opens in new tab)
Share

Frontier multimodal medical AI systems are primarily evaluated on Western benchmarks featuring pristine, high-resolution pathology scans. However, clinical deployment in the Global South relies on ad-hoc "telepathology", capturing microscope views via budget smartphones and transmitting them over heavily compressed networks like WhatsApp. We introduce GSM-PathEval, a targeted benchmark that programmatically applies real-world infrastructural degradations (eyepiece glare, lens blur, and heavy JPEG quantization) to the gold-standard PathVQA dataset. Using Adaption Labs, we demonstrate a critical "Reality Gap": frontier Vision-Language Models optimized for clean imagery suffer drastic diagnostic accuracy drops delta(A) under these simulated field conditions. This performance cliff poses a severe misdiagnosis risk, proving that standard capabilities benchmarks are insufficient for safe medical AI deployment in low-resource environments.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project addresses an important deployment challenge that is often overlooked in medical AI: how vision-language models perform under the real-world infrastructure constraints common in the Global South. The degradation pipeline is practical, reproducible, and highlights meaningful safety risks that traditional benchmarks fail to capture. To strengthen the work, I would encourage expanding the evaluation beyond the current sample size and incorporating a wider range of realistic imaging conditions, devices, and multilingual clinical settings. Comparing additional models and exploring mitigation strategies would also increase the practical impact. Overall, this is a thoughtful and well-focused benchmark with strong real-world relevance.

  2. Topics seems very practical. It'd be great to bring more clarity on which AI you tested and see if results hold in more than one evaluation setup. More tests across other models would make the results stronger.

  3. Focused benchmark with a useful reframe: stress medical VLMs with realistic infrastructure degradation (glare, blur, WhatsApp-grade compression) instead of synthetic adversarial noise. The threat model is crisp and the pipeline is reproducible from a public dataset. The limits are scale and specification: 50 binary questions, one fixed degradation profile, no significance test, and the evaluated model is never named, which makes the 84% to 52% drop hard to interpret or reproduce. The degradation parameters are hand-picked rather than calibrated against real telepathology images, and the visual-evidence section describes failures without showing any. Name the model, grow the sample, report variance, justify the parameters against real device and network measurements, and include paired pristine-vs-degraded examples where the prediction flips.

Cite this project

@misc{awasthi2026gsmpatheval,
  title = {{GSM-PathEval: A Global South Robustness Benchmark for Telepathology Vision-Language Models}},
  author = {Anshuman Awasthi},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/gsmpatheval-a-global-south-robustness-benchmark-for-telepathology-visionlanguage-models-1al7}},
  url = {https://apartresearch.com/sprints/projects/gsmpatheval-a-global-south-robustness-benchmark-for-telepathology-visionlanguage-models-1al7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026