PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud

Patrick Passos

PixTrap is a Brazilian Portuguese safety benchmark evaluating whether LLMs refuse Pix fraud and social-engineering misuse while still answering legitimate anti-fraud requests. English-centric evaluations miss regional idioms, local institutions, and Brazil-specific scam patterns. By pairing harmful prompts with benign near-neighbors, PixTrap measures both unsafe compliance and over-refusal. Across five models in pt-BR and English, we find a modest 10-20% cross-language safety gap. PixTrap is a reproducible package and a reusable recipe for local-safety benchmarks in underrepresented regions.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This matters, and it's about to matter more. Pix is already the rail 150M+ Brazilians run on, and as other Latam countries copy the model (Colombia's Bre-B is the obvious next one), the fraud patterns travel with it. A safety benchmark grounded in the actual scams, in the actual language, is the right thing to be building now, not after the copies ship. The near-neighbor design is the smart call: instead of just asking whether a model refuses fraud, you pair each scam prompt with a legit look-alike and check whether it can still help the person writing a bank-security article. That's the question that matters once it's deployed. Safe redirect is the part I hadn't seen framed this way. A flat "I can't help with that" is dead weight to someone who's mid-scam, and putting Llama at 0% next to Kimi at 100% makes that obvious. But the thing I'd credit most is the scorer bug. Your first pass showed a 50-70% cross-language gap. You dug in, found it was "I cannot" matching while the models wrote "I can't," fixed it, and the gap fell to 0-20%. Most people ship the scary number and move on. You caught yourself.

The place to push hardest: your headline gap still leans on the same scorer that already burned you once. You say it yourself. The keyword matcher agrees with your own manual labels only 73% of the time, and the 10-20% gap that survived the fix sits inside intervals as wide as 6% to 51% at ten prompts a cell. Three moves. One, re-score with an LLM judge, or run both, so the cross-language claim isn't coming from the tool that inverted last time. Fixing one contraction doesn't prove there's no second pattern you haven't tripped yet. Two, validate the scorer language by language against manual labels, not 30 samples pooled together. Your own lesson is that scorers don't carry across languages, so one blended 73% can't tell you whether Portuguese and English agree at the same rate, and that rate is the whole gap. Three, either grow the prompt set so the intervals can actually hold a 10-20% gap, or pull the gap out of the headline and lead with safe redirect. 0% versus 100% is signal ten prompts already support.

Solid work. What excites me is where it goes next: the same recipe pointed at Bre-B, or any central bank shipping an instant-payment rail. Build that one and I'll read it too.

The most urgent improvement is replacing keyword-based scoring with an LLM-as-judge for at least a subset of responses, as the paper itself recommends. At 73% agreement, keyword scoring is too noisy to support the per-model comparisons the paper makes. The paper is admirably honest about this, but the honesty comes at the cost of weakening the empirical claims. Even scoring 50 samples manually with a second native-speaker annotator could provide a more reliable ground truth than the single-author delayed audit.

The reusability claim, that PixTrap is a recipe for other regions (UPI in India, M-Pesa in Kenya), is interesting but is one sentence. Expanding this slightly, with a brief description of what would need to change (fraud taxonomy, language, payment institution names) and what stays constant (near-neighbor design, calibration score, scorer validation approach), would make the contribution more concrete and help others actually apply it.

use just portuguese next time

Very interesting project and the contributions to the literature are clearly laid out. What it would take to make these results more impactful is also made explicit to help determine the validity of the results. The replicability for other payment systems with similar fraud patterns is also persuasive when it comes to impact.

Cite this work

@misc {

title={

(HckPrj) PixTrap Reveals LLM Safety Calibration Gaps in Brazilian Pix Fraud

},

author={

Patrick Passos

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.