MediShield-Proxy: A Local Privacy-Preserving Intermediary Layer for Secure Clinical LLM Ingestion
Genius Tanaka Chipfupa
MediShield -Proxy is a lightweight, zero-trust Local Area Network (LAN) middleware architecture designed for African medical institutions. It intercepts sensitive patient data locally, automatically masking demographics, histories, and clinical conditions with reversible tokens before they can be leaked to external cloud-based Large Language Models (LLMs). This tool allows resource-constrained facilities to safely leverage global AI diagnostic and administrative utility while strictly maintaining data sovereignty and regional data protection compliance.
This goes after a real problem that doesn't get enough attention clinicians in low-resource settings pasting identifiable patient data straight into ChatGPT or Claude — and the basic design instinct is right. Reversible local pseudonymization, mapping kept only in memory, re-identification on the way back: that's the correct shape given the constraint, and keeping it regex-based and CPU-only makes sense for old clinic hardware. The part I valued most isn't the tool, it's the insight behind it — that stripping direct identifiers isn't enough, because a rare pathology plus a local landmark can re-identify someone on its own. That's true and most people miss it, and abstracting geography into epidemiologically-equivalent tiers is a smart way to handle it. I also want to credit the negative results section; reporting that the embedding filter added 4,200ms and tagged "breakbone fever" as a location is exactly the kind of honesty I want to see.
Where it falls down is the evaluation. The 97.1% F1 — and especially the perfect 100% on National IDs — comes from 150 synthetic notes the team wrote themselves, so it's really measuring how well the regex matches the cases the author already had in mind, not messy real clinical text. That 100% should worry you, not reassure you; it's a sign the test set is circular. The lowercase-name and code-switching failures you flag are exactly the things that'll be far more common in real notes than in your synthetic ones. I'd also push back on the framing: you describe this as a zero-trust LAN interceptor that blocks outbound packets at the gateway, but what you've actually built is a client-side proxy the user has to choose to route through — nothing stops a clinician just opening chatgpt.com directly. That's a different threat model, and it should be said plainly. Three things would help a lot: test it on a clinical corpus you didn't write yourself, dial the enforcement claims back to what the proxy really guarantees, and deal with the obvious leak you don't address the raw symptoms and rare conditions still go to the external model unmasked, which in a small community can identify someone by itself. The dual-use note on the inversion module was a good call to include.
Healthcare systems in resource-constrained settings must protect patient privacy while accessing the capabilities of external large language models. This paper addresses that balancing act with a practical and well-motivated solution. The local intermediary architecture stands out as the submission's strongest contribution, proposing a path to reducing privacy exposure without cutting off access to advanced AI tools. The evaluation is the area most in need of development. The current reliance on a small synthetic dataset limits the generalisability of the findings, and the work would benefit from benchmarking against established de-identification methods and validation on more representative clinical data. Engaging more directly with the system's limitations, particularly its performance in multilingual or adversarial settings, would give the conclusions greater credibility and extend the paper's contribution to the broader AI safety literature.
1. PII to AI is a real problem - there are many approaches but I don't believe this approach solves for scale, given the false positives and latency
2. Consider how tokenization can be solved not "at rest" but while in motion
Cite this work
@misc {
title={
(HckPrj) MediShield-Proxy: A Local Privacy-Preserving Intermediary Layer for Secure Clinical LLM Ingestion
},
author={
Genius Tanaka Chipfupa
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


