PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud
Patrick Passos
PixTrap is a Brazilian Portuguese safety benchmark evaluating whether LLMs refuse Pix fraud and social-engineering misuse while still answering legitimate anti-fraud requests. English-centric evaluations miss regional idioms, local institutions, and Brazil-specific scam patterns. By pairing harmful prompts with benign near-neighbors, PixTrap measures both unsafe compliance and over-refusal. Across five models in pt-BR and English, we find a modest 10-20% cross-language safety gap. PixTrap is a reproducible package and a reusable recipe for local-safety benchmarks in underrepresented regions.
This matters, and it's about to matter more. Pix is already the rail 150M+ Brazilians run on, and as other Latam countries copy the model (Colombia's Bre-B is the obvious next one), the fraud patterns travel with it. A safety benchmark grounded in the actual scams, in the actual language, is the right thing to be building now, not after the copies ship. The near-neighbor design is the smart call: instead of just asking whether a model refuses fraud, you pair each scam prompt with a legit look-alike and check whether it can still help the person writing a bank-security article. That's the question that matters once it's deployed. Safe redirect is the part I hadn't seen framed this way. A flat "I can't help with that" is dead weight to someone who's mid-scam, and putting Llama at 0% next to Kimi at 100% makes that obvious. But the thing I'd credit most is the scorer bug. Your first pass showed a 50-70% cross-language gap. You dug in, found it was "I cannot" matching while the models wrote "I can't," fixed it, and the gap fell to 0-20%. Most people ship the scary number and move on. You caught yourself.
The place to push hardest: your headline gap still leans on the same scorer that already burned you once. You say it yourself. The keyword matcher agrees with your own manual labels only 73% of the time, and the 10-20% gap that survived the fix sits inside intervals as wide as 6% to 51% at ten prompts a cell. Three moves. One, re-score with an LLM judge, or run both, so the cross-language claim isn't coming from the tool that inverted last time. Fixing one contraction doesn't prove there's no second pattern you haven't tripped yet. Two, validate the scorer language by language against manual labels, not 30 samples pooled together. Your own lesson is that scorers don't carry across languages, so one blended 73% can't tell you whether Portuguese and English agree at the same rate, and that rate is the whole gap. Three, either grow the prompt set so the intervals can actually hold a 10-20% gap, or pull the gap out of the headline and lead with safe redirect. 0% versus 100% is signal ten prompts already support.
Solid work. What excites me is where it goes next: the same recipe pointed at Bre-B, or any central bank shipping an instant-payment rail. Build that one and I'll read it too.
The most urgent improvement is replacing keyword-based scoring with an LLM-as-judge for at least a subset of responses, as the paper itself recommends. At 73% agreement, keyword scoring is too noisy to support the per-model comparisons the paper makes. The paper is admirably honest about this, but the honesty comes at the cost of weakening the empirical claims. Even scoring 50 samples manually with a second native-speaker annotator could provide a more reliable ground truth than the single-author delayed audit.
The reusability claim, that PixTrap is a recipe for other regions (UPI in India, M-Pesa in Kenya), is interesting but is one sentence. Expanding this slightly, with a brief description of what would need to change (fraud taxonomy, language, payment institution names) and what stays constant (near-neighbor design, calibration score, scorer validation approach), would make the contribution more concrete and help others actually apply it.
use just portuguese next time
Very interesting project and the contributions to the literature are clearly laid out. What it would take to make these results more impactful is also made explicit to help determine the validity of the results. The replicability for other payment systems with similar fraud patterns is also persuasive when it comes to impact.
Cite this work
@misc {
title={
(HckPrj) PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud
},
author={
Patrick Passos
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


