Skip to content
Sprint projectJun 21, 2026Harare, Zimbabwe

MediShield-Proxy: A Local Privacy-Preserving Intermediary Layer for Secure Clinical LLM Ingestion

Genius Tanaka Chipfupa · Team MediShield

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: MediShield-Proxy: A Local Privacy-Preserving Intermediary Layer for Secure Clinical LLM Ingestion

Recording (opens in new tab)Code (opens in new tab)
Share

MediShield -Proxy is a lightweight, zero-trust Local Area Network (LAN) middleware architecture designed for African medical institutions. It intercepts sensitive patient data locally, automatically masking demographics, histories, and clinical conditions with reversible tokens before they can be leaked to external cloud-based Large Language Models (LLMs). This tool allows resource-constrained facilities to safely leverage global AI diagnostic and administrative utility while strictly maintaining data sovereignty and regional data protection compliance.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This goes after a real problem that doesn't get enough attention clinicians in low-resource settings pasting identifiable patient data straight into ChatGPT or Claude — and the basic design instinct is right. Reversible local pseudonymization, mapping kept only in memory, re-identification on the way back: that's the correct shape given the constraint, and keeping it regex-based and CPU-only makes sense for old clinic hardware. The part I valued most isn't the tool, it's the insight behind it — that stripping direct identifiers isn't enough, because a rare pathology plus a local landmark can re-identify someone on its own. That's true and most people miss it, and abstracting geography into epidemiologically-equivalent tiers is a smart way to handle it. I also want to credit the negative results section; reporting that the embedding filter added 4,200ms and tagged "breakbone fever" as a location is exactly the kind of honesty I want to see.

    Where it falls down is the evaluation. The 97.1% F1 — and especially the perfect 100% on National IDs — comes from 150 synthetic notes the team wrote themselves, so it's really measuring how well the regex matches the cases the author already had in mind, not messy real clinical text. That 100% should worry you, not reassure you; it's a sign the test set is circular. The lowercase-name and code-switching failures you flag are exactly the things that'll be far more common in real notes than in your synthetic ones. I'd also push back on the framing: you describe this as a zero-trust LAN interceptor that blocks outbound packets at the gateway, but what you've actually built is a client-side proxy the user has to choose to route through — nothing stops a clinician just opening chatgpt.com directly. That's a different threat model, and it should be said plainly. Three things would help a lot: test it on a clinical corpus you didn't write yourself, dial the enforcement claims back to what the proxy really guarantees, and deal with the obvious leak you don't address the raw symptoms and rare conditions still go to the external model unmasked, which in a small community can identify someone by itself. The dual-use note on the inversion module was a good call to include.

    Read full reviewShow less
  2. Healthcare systems in resource-constrained settings must protect patient privacy while accessing the capabilities of external large language models. This paper addresses that balancing act with a practical and well-motivated solution. The local intermediary architecture stands out as the submission's strongest contribution, proposing a path to reducing privacy exposure without cutting off access to advanced AI tools. The evaluation is the area most in need of development. The current reliance on a small synthetic dataset limits the generalisability of the findings, and the work would benefit from benchmarking against established de-identification methods and validation on more representative clinical data. Engaging more directly with the system's limitations, particularly its performance in multilingual or adversarial settings, would give the conclusions greater credibility and extend the paper's contribution to the broader AI safety literature.

    Read full reviewShow less
  3. 1. PII to AI is a real problem - there are many approaches but I don't believe this approach solves for scale, given the false positives and latency

    2. Consider how tokenization can be solved not "at rest" but while in motion

Cite this project

@misc{chipfupa2026medishieldproxy,
  title = {{MediShield-Proxy: A Local Privacy-Preserving Intermediary Layer for Secure Clinical LLM Ingestion}},
  author = {Genius Tanaka Chipfupa},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/medishieldproxy-a-local-privacypreserving-intermediary-layer-for-secure-clinical-llm-ingestion-2sp5}},
  url = {https://apartresearch.com/sprints/projects/medishieldproxy-a-local-privacypreserving-intermediary-layer-for-secure-clinical-llm-ingestion-2sp5}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026