AyuGuard: A Safety-Routing Framework and Evaluation Benchmark for Localized Pharmacology in Indian Rural Healthcare.

Rewant Anand

AyuGuard is a deterministic middleware safety-routing framework and offline edge-triage system designed to mitigate life-threatening AI hallucinations in rural Indian healthcare. It specifically targets the "epistemological asymmetry" where frontier LLMs confidently hallucinate safe outcomes for dangerous interactions between allopathic pharmaceutical drugs and traditional AYUSH remedies. By intercepting high-risk queries using an ultra-low-latency (+112ms) pattern-matching router, AyuGuard blocks harmful drug-herb combinations before they ever reach the primary LLM, reducing dangerous hallucinations by 93%. Additionally, the platform equips frontline ASHA workers with an offline-first NEWS2 physiological triage engine featuring Hinglish voice-dictation, ensuring safe, localized clinical assessments remain functional even in environments with intermittent 2G/3G connectivity

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

AyuGuard is a highly relevant AI safety project because it targets a concrete, high-risk failure mode: LLMs giving confident but unsafe advice about allopathic medicine and AYUSH/traditional remedies in rural Indian healthcare contexts. The project does a strong job explaining why this is not just a general medical QA issue, but a localization and safety problem where model hallucinations could plausibly lead to severe harm. The reported improvement including false negatives, false positives, F1 score, latency, and manual error analysis from AyuGuard is also compelling. The App also shows concrete examples and visualizations of answers and LLM filtering for practical usages.

One area for improvement is the dataset and language coverage. The domain-context queries from the .js/.tsx file appear to cover around 150 queries, but stronger evaluation would likely require broader Hindi and Hinglish coverage, especially for symptom descriptions, treatment terms, AYUSH remedies, and common rural phrasing. It would also be useful to analyze which types of queries still produce false positives and false negatives, and whether adding more context-specific language examples reduces these errors.

I would also like to see the evaluation broken down by language: Hindi, Hinglish, and English. This would make the safety alignment claim stronger because it would show whether AyuGuard performs consistently across different language contexts, rather than only improving aggregate scores.

In addition, the dataset includes patient risk levels for each prompt. It would be valuable to report how true positives, false positives, false negatives, and F1 scores change across different risk levels, from high-risk prompts to low-risk prompts.

Overall, AyuGuard is one of the more practically safety-relevant projects because it connects LLM localization failures to a specific harm pathway. It would be even stronger with clearer benchmark extensions and specific language/risk level analysis.

A safety filter that sits in front of an AI assistant for rural health workers and blocks dangerous advice about mixing AYUSH remedies with prescription drugs. Good problem, real user, and you shipped a working prototype with offline triage, which shows you built for the actual setting. Deferring to a clinician instead of answering is the right call. Since the filter is the whole product, everything rests on how reliably it catches the dangerous cases, and that is where I'd want more proof. It matches keywords, so phrasing it hasn't seen can slip through, and in your own demo a disguised query was judged low risk. Your test questions are all plainly worded, so I'd read the 96.5% as a catch rate on easy inputs, not the messy phrasing a real health worker would use. Next step I'd prioritise: test paraphrased and disguised versions of the same dangerous questions and report how many it still catches.

AyuGuard identifies a highly relevant and high-impact AI safety problem in rural healthcare and proposes a practical mitigation strategy. It pins down a specific, high-stakes failure (models waving through allopathic and AYUSH combinations for automation-biased ASHA workers) and evaluates a fix end to end, 42% to 96.5% true-positive rate with false-positive rate and latency reported. The case for deterministic middleware over model self-alignment is well argued, and the error analysis is the highlight: three traceable failures with root causes and concrete fixes. The catch is construct validity. A regex router that fires when a Schedule-H term and an AYUSH term co-occur is being tested on a benchmark built around exactly those co-occurrences, so part of the headline number reflects the test matching the matcher's design. The 150-query set is author-built with no external clinical-label check, and the NEWS2 engine is included but never evaluated. Name the circularity directly, add held-out phrasings the matrix wasn't designed for, get a pharmacologist to validate a sample of the labels, and either evaluate NEWS2 or scope it out.

Cite this work

@misc {

title={

(HckPrj) AyuGuard: A Safety-Routing Framework and Evaluation Benchmark for Localized Pharmacology in Indian Rural Healthcare.

},

author={

Rewant Anand

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.