Skip to content
Sprint projectJun 21, 2026Hồ Chí Minh

V-Sentinel

Lê Huy Hùng, Nguyễn Trần Minh Tâm · Team 2vane

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

AI safety tooling is overwhelmingly built and benchmarked in English and degrades sharply on low-resource languages — a critical gap as Vietnam deploys AI in public services, healthcare, and education under Decree 142/2026/ND-CP. We present V-Sentinel, a dual-control guardrail that pairs a deterministic, OWASP-tagged rule backbone with an LLM safety classifier and fails closed. A non-destructive normalization layer neutralizes Vietnamese-specific evasion (teencode, leetspeak, diacritic folding, Base64) without corrupting legitimate text; a domain-aware GraphRAG layer attaches the jurisdictionally correct legal frame (ND-142/PDPD, FERPA, COPPA, HIPAA); a REFRAME mechanism converts over-refusal of legitimate sensitive queries into responsible, cited answers; and an output stage redacts Vietnamese PII. On two Vietnamese benchmarks, V-Sentinel flags ∼90% of harmful prompts and the rule backbone alone blocks 64% of injection attacks with no model call — still flagging 100% under classifier outage. An ablation shows the two controls are complementary rather than redundant: rules carry the obfuscation/injection axis and provide the fail-closed guarantee, while the classifier carries content-harm. Against published Llama Guard results, which fall to 51% recall on adversarial prompts and degrade multilingually, V-Sentinel sustains ∼90% on Vietnamese. The main takeaway: a thin, local-first, law-grounded guardrail makes an open model compliant without the cost of fine-tuning a sovereign model, and re-targets to any jurisdiction by swapping its rule pack and legal graph.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. V Sentinel addresses a genuine gap: English centric guardrails degrade on Vietnamese, and the populations most reliant on public sector AI are the least protected. The dual control architecture, pairing a deterministic rule backbone with an LLM classifier, is a well motivated design choice. The strongest evidence for this is the ablation in Table 3, which shows the two controls cover complementary axes. Rules handle injection and obfuscation; the classifier handles content harm. Neither alone spans the full threat surface, and their combination provides measurable fail closed coverage under classifier outage. This is the paper's most convincing result. The legal grounding via GraphRAG is an interesting addition that ties decisions to specific articles of Vietnamese law, though its validation in the paper is light. The over refusal story needs more scrutiny. The paper presents REFRAME as a solution to over refusal, but 11% hard refusal plus 54% of benign prompts receiving a cautioned answer means 65% of legitimate queries get a degraded experience. For a citizen interacting with a public service chatbot, receiving a hedged, citation laden response to a straightforward health or education question is still a usability cost. A breakdown of which benign query types trigger REFRAME, and whether users perceive it as helpful or obstructive, would strengthen the argument that this is mitigation rather than a different flavor of the same problem.

    Evaluation rigor is another area for improvement. Results are single run on two benchmarks without confidence intervals. The deterministic rule results reproduce by design, which is a genuine strength, but the classifier dependent numbers could vary across runs and this is never quantified. Reporting even minimal variance (three runs, standard deviation) would add credibility at low cost.

    The scope is appropriate for a hackathon. The system is open source, the test suite is substantial (151 tests across 27 files), and the architecture is clearly described. The writing is generally clear and well structured. The main areas to develop post hackathon are: running direct baselines on Vietnamese prompts, building a dedicated over refusal suite, and addressing the subtle harm categories (disinformation, discrimination) where detection currently sits at 20 to 58 percent.

    Read full reviewShow less
  2. A well-executed systems-oriented AI safety project addressing an important gap in multilingual guardrails for Vietnamese public-sector applications. The dual-control architecture combining deterministic rules with an LLM classifier is well motivated, and the integration of legal grounding, PII protection, and fail-closed behavior demonstrates strong practical engineering.

    The project would be strengthened by:

    - Expanding evaluation to adaptive jailbreaks, code-switching attacks, and additional robustness benchmarks.

    - Reporting confidence intervals and latency measurements to better demonstrate deployment readiness.

    - Providing deeper analysis of false positives and over-refusal

  3. Amazing concept. Keep up the good work.

  4. Good project idea which solve the existed gap in using AI for everyone in Vietnam

Cite this project

@misc{hung2026vsentinel,
  title = {{V-Sentinel}},
  author = {Lê Huy Hùng and Nguyễn Trần Minh Tâm},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/vsentinel-bio7}},
  url = {https://apartresearch.com/sprints/projects/vsentinel-bio7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026