Skip to content
Sprint projectJun 22, 2026Brazil

PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud

Patrick Passos · Team PixTrap

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud

Presentation

Presentation: PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud

Code (opens in new tab)
Share

PixTrap is a Brazilian Portuguese safety benchmark evaluating whether LLMs refuse Pix fraud and social-engineering misuse while still answering legitimate anti-fraud requests. English-centric evaluations miss regional idioms, local institutions, and Brazil-specific scam patterns. By pairing harmful prompts with benign near-neighbors, PixTrap measures both unsafe compliance and over-refusal. Across five models in pt-BR and English, we find a modest 10-20% cross-language safety gap. PixTrap is a reproducible package and a reusable recipe for local-safety benchmarks in underrepresented regions.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The most urgent improvement is replacing keyword-based scoring with an LLM-as-judge for at least a subset of responses, as the paper itself recommends. At 73% agreement, keyword scoring is too noisy to support the per-model comparisons the paper makes. The paper is admirably honest about this, but the honesty comes at the cost of weakening the empirical claims. Even scoring 50 samples manually with a second native-speaker annotator could provide a more reliable ground truth than the single-author delayed audit.

    The reusability claim, that PixTrap is a recipe for other regions (UPI in India, M-Pesa in Kenya), is interesting but is one sentence. Expanding this slightly, with a brief description of what would need to change (fraud taxonomy, language, payment institution names) and what stays constant (near-neighbor design, calibration score, scorer validation approach), would make the contribution more concrete and help others actually apply it.

    Read full reviewShow less
  2. This matters, and it's about to matter more. Pix is already the rail 150M+ Brazilians run on, and as other Latam countries copy the model (Colombia's Bre-B is the obvious next one), the fraud patterns travel with it. A safety benchmark grounded in the actual scams, in the actual language, is the right thing to be building now, not after the copies ship. The near-neighbor design is the smart call: instead of just asking whether a model refuses fraud, you pair each scam prompt with a legit look-alike and check whether it can still help the person writing a bank-security article. That's the question that matters once it's deployed. Safe redirect is the part I hadn't seen framed this way. A flat "I can't help with that" is dead weight to someone who's mid-scam, and putting Llama at 0% next to Kimi at 100% makes that obvious. But the thing I'd credit most is the scorer bug. Your first pass showed a 50-70% cross-language gap. You dug in, found it was "I cannot" matching while the models wrote "I can't," fixed it, and the gap fell to 0-20%. Most people ship the scary number and move on. You caught yourself.

    The place to push hardest: your headline gap still leans on the same scorer that already burned you once. You say it yourself. The keyword matcher agrees with your own manual labels only 73% of the time, and the 10-20% gap that survived the fix sits inside intervals as wide as 6% to 51% at ten prompts a cell. Three moves. One, re-score with an LLM judge, or run both, so the cross-language claim isn't coming from the tool that inverted last time. Fixing one contraction doesn't prove there's no second pattern you haven't tripped yet. Two, validate the scorer language by language against manual labels, not 30 samples pooled together. Your own lesson is that scorers don't carry across languages, so one blended 73% can't tell you whether Portuguese and English agree at the same rate, and that rate is the whole gap. Three, either grow the prompt set so the intervals can actually hold a 10-20% gap, or pull the gap out of the headline and lead with safe redirect. 0% versus 100% is signal ten prompts already support.

    Solid work. What excites me is where it goes next: the same recipe pointed at Bre-B, or any central bank shipping an instant-payment rail. Build that one and I'll read it too.

    Read full reviewShow less
  3. use just portuguese next time

  4. Very interesting project and the contributions to the literature are clearly laid out. What it would take to make these results more impactful is also made explicit to help determine the validity of the results. The replicability for other payment systems with similar fraud patterns is also persuasive when it comes to impact.

Cite this project

@misc{passos2026pixtrap,
  title = {{PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud}},
  author = {Patrick Passos},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/pixtrap-reveals-llm-safety-calibration-gaps-in-brazilian-pix-fraud-c2a2}},
  url = {https://apartresearch.com/sprints/projects/pixtrap-reveals-llm-safety-calibration-gaps-in-brazilian-pix-fraud-c2a2}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026