Skip to content
Sprint projectJun 22, 2026Ho Chi Minh City

Pragmatic Sophistry in Vietnamese Multi-Agent Oversight

Bui Cao Dong Nghi, Truong Ngoc Minh Tam, Bui Ngoc Linh · Team AI Safety Enthusiasts

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow agents in which control violations are expressed not through explicit risk keywords but through pragmatic, culturally fluent language (hierarchy, relationship pressure, euphemism, omission, and routine framing). We build a four-stage pipeline (attacker-renderer, fluent oracle, naive executor, monitor) and evaluate two LLM monitors (Gemini 2.5 Flash, DeepSeek Chat) under three visibility conditions (seeing the original instruction, only the executor's action/report, or the full instruction-action trace) across three domains: e-commerce, banking, and government. Our central finding is a context-compression failure that is strongly domain-dependent. In e-commerce, monitoring only the executor's action raised false negatives on suspicious pragmatic-Vietnamese requests from 6.4% to 22.3%, with full-trace monitoring recovering performance to 8.5%; up to 24.4% of violations catchable from the instruction were missed once only the action was shown. This failure was mild in banking and absent in government, where violations are concrete, explicit actions that survive compression. Across all domains, the rare misses were high-confidence (0.90–1.00). We conclude that the executor/report layer can disguise safety-relevant context in some domains, that the vulnerability tracks the type of violation rather than language alone, and that preserving the instruction-action trace is an important, low-cost safeguard for multilingual agentic oversight.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The paper does a good job showing how safety relevant context can get lost when a multilingual instruction is compressed into a downstream action or report. I especially liked the instruction -> action > trace comparison, as it makes the core finding practical and easy to understand.

    The results are meaningful, particularly the e-commerce case where action only monitoring clearly underperforms and full trace monitoring recovers much of the lost signal. Adding human generated context and inputs would have been very helpful. That would make the results feel more grounded and help confirm whether the same patterns appear outside model generated workflows.

  2. Empirically grounded — Concrete numbers (6.4% → 22.3% false negatives, recovery to 8.5%) make the findings actionable

    Domain comparison — Testing across e-commerce, banking, and government reveals that the vulnerability is not uniform, which is a nuanced and useful finding

    Practical recommendation — "Preserve the instruction-action trace" is a low-cost, implementable safeguard

    Timely — Directly relevant to the current wave of agentic AI deployment

    This is a publishable, original contribution that addresses a real gap identified by major safety reports and aligns with where the top labs (Anthropic, DeepMind) are investing right now. It would be a strong fit for venues like EMNLP, ACL, AAAI (safety track), or workshops like SafeGenAI/TrustNLP.

  3. Your discovery that an AI assistant's routine summarisation alone — no deception needed — is enough to hide a safety violation from a monitor is a genuinely eye-opening finding, and mapping exactly which domains this is dangerous in makes it immediately useful for real teams. To push it further, test with an even more aggressive summariser to show the worst-case risk, and consider releasing your Vietnamese test scenarios as a public benchmark — nothing like it exists today, and the field would build on it for years.

  4. This is a strong and practically motivated project. The paper identifies a clear oversight failure mode: safety-relevant context can disappear when a Vietnamese instruction is compressed into an executor action/report, especially in relational or personally identifiable information workflows. The instruction/action/trace split is a good experimental design choice, and the cross-domain comparison makes the central claim more interesting than a generic “multilingual monitoring is hard” result.

    The strongest contribution is the domain-dependent mechanism: e-commerce failures are not just caused by Vietnamese pragmatics, but by executor-side context compression of authorization, relationship, and omission cues. The worked examples make this intuitive, and the conditional report-hiding metric is useful.

    The main improvement would be to strengthen the empirical grounding. The datasets are small, model-generated, and filtered by model/oracle components, so the conclusions should be framed more clearly as an audit prototype rather than broad evidence about Vietnamese workflows. The paper would benefit from more detail on scenario generation, oracle criteria, executor prompts, monitor prompts, and exact labeling rules so others can reproduce the results. It would also help to include human Vietnamese annotators or at least a small human validation sample, since the core claim depends on culturally fluent pragmatic meaning being recoverable.

    The statistical presentation is promising but could be cleaner. Some claims rely on small cells, especially cue/persona/pressure breakdowns, and should be marked more explicitly as exploratory. Government results are described as “absent,” but the paper correctly notes the study is underpowered there; I would consistently phrase this as “no measurable gap in this sample” rather than “absent.” Finally, the paper should separate more sharply what is caused by multilingual pragmatics, what is caused by executor summarization, and what may be caused by domain-specific task structure.

    Overall, this is a useful and novel AI safety audit with a clear practical takeaway: action-only monitoring can be unsafe in domains where violations depend on social/authorization context, and preserving the instruction-action trace is a low-cost mitigation.

    Read full reviewShow less

Cite this project

@misc{nghi2026pragmatic,
  title = {{Pragmatic Sophistry in Vietnamese Multi-Agent Oversight}},
  author = {Bui Cao Dong Nghi and Truong Ngoc Minh Tam and Bui Ngoc Linh},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/pragmatic-sophistry-in-vietnamese-multiagent-oversight-u5xy}},
  url = {https://apartresearch.com/sprints/projects/pragmatic-sophistry-in-vietnamese-multiagent-oversight-u5xy}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026