Skip to content
Sprint projectJun 22, 2026Ho Chi Minh City, Vietnam

Vietnamese RAG Prompt Injection Test Kit

Hoàng Trọng Trà · Team SYP

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Vietnamese RAG Prompt Injection Test Kit

Presentation

Presentation: Vietnamese RAG Prompt Injection Test Kit

Code (opens in new tab)
Share

A Vietnamese-language safety benchmark and evaluation toolkit for document-level prompt injection in retrieval-augmented generation (RAG) systems. The project provides 48 synthetic test cases across six practical domains, a reproducible evaluation harness, deterministic and real-model testing, manual-review exports, and lightweight mitigation controls. It helps Vietnamese and Southeast Asian RAG builders test whether retrieved documents can override system intent, leak sensitive data, hijack citations, or produce unsafe advice before deployment.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I think this project is highly relevant to AI safety because Vietnamese-language RAG prompt-injection benchmarks are still scarce, while RAG chatbots are increasingly used across HR, customer support, banking/fintech, education, healthcare, and public-service settings in Vietnam. The project makes a strong contribution by introducing a defensive benchmark for Vietnamese RAG systems and comparing baseline behavior against strategies such as instruction-hierarchy prompting and context sanitization.

    The methodology is generally sound. The report clearly presents the benchmark setup, attack categories, mitigation controls, automatic scoring, and manual review. The results also communicate a meaningful safety finding: the baseline system showed nontrivial attack success, while the mitigation controls reduced failures in the reviewed cases.

    However, the report could explain the findings more clearly. For example, IDs such as TC014, TC027, TC046, etc. are referenced in the analysis, but their meanings and data structures are not clearly explained for readers. It would help to briefly describe what each highlighted test case represents and why it matters. The report seems to emphasize Table 1, Figure 1, and Table 2, but Table 3 and Table 4 could be discussed in more depth. In particular, the authors could explain what the results suggest about Vietnamese-only attacks versus Vietnamese-English code-switching attacks, as well as how the different GPT-5.4-mini evaluation settings compare across the reported metrics.

    Another useful addition would be a sector-level analysis. Since the benchmark covers HR, customer support, banking/fintech, school policy, clinic information, and public-service FAQ domains, the authors could discuss whether certain sectors appear more vulnerable or safety-critical. Even with only 48 test cases, highlighting qualitative patterns across sectors would strengthen the impact of the report.

    Overall, this is a good and practically useful project. It provides a meaningful starting point for safer Vietnamese and Southeast Asian RAG deployment, and the report would be even stronger with clearer interpretation of tables, domain-specific risks, and GPT-5.4 model variations.

    Read full reviewShow less
  2. This is the kind of thing regional teams can actually run before launch, which is the point. Your evaluation habits are the strongest part: the manual pass caught both false positives and four real failures the auto-judge missed. The code-switching result (87.5% at baseline) is worth leading with. One quick fix: the submitted PDF still has the template's placeholder text sitting in the Results and References sections, plus a "Miltigated" typo. After that, headline the real-model number (10.42%) and keep it clear of the deterministic sim, whose 54% comes from a stub built to fail. Then push past 48 cases and add a second model. Good, practical safety work.

  3. Good job! You should add held-out validation, as you tested your defenses on the same 48 attacks that you designed the defenses on. In addition, both defenses scored a perfect 0/48, so there is no signal to distinguish the two controls. Consider tests that are hard enough to separate them.

    Also remember to remove the leftover Apart template boilerplate in Section 4 ("Present your main findings…").

Cite this project

@misc{tra2026vietnamese,
  title = {{Vietnamese RAG Prompt Injection Test Kit}},
  author = {Hoàng Trọng Trà},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/vietnamese-rag-prompt-injection-test-kit-vrju}},
  url = {https://apartresearch.com/sprints/projects/vietnamese-rag-prompt-injection-test-kit-vrju}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026