Los peajes de los de abajo

Mongui Rogers

Los LLM ya resisten la psicofancia clásica en español: no validan el dato falso. Pero cobran un peaje distinto cuando el hablante usa jerga regional colombiana — no preguntan, fingen comprensión y malinterpretan términos de alto riesgo (leyeron "vacuna" como droga, cuando significa extorsión). Medimos tres caras del peaje (tokens, calidad, comprensión) con 10 probes validados por nativo y doble juez (LLM + humano). La salvaguarda aguanta; la equidad de comprensión no. El juez-LLM sub-cuenta el daño en lenguas no dominantes: se necesita un evaluador nativo. El peaje cae sobre el habla más local.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This project highlights an important and often overlooked AI safety issue: the unequal ability of LLMs to understand regional language varieties. The methodology is thoughtful, and the distinction between classical sycophancy and “comprehension sycophancy” is particularly compelling. The finding that LLM judges can underestimate these failures adds further value and practical relevance. Expanding the dataset and adding more annotators would strengthen the results, but the work already makes a meaningful contribution to multilingual AI safety.

You're asking a good question. If the safeguard against explicit error already holds up in Spanish, where does the failure move to once high-stakes regional slang comes in? That's the right way to split the problem apart, and it's the strongest part of the project, you're not just asking whether the model holds up against an obvious error, but whether it actually understands what's being said.

One area where you could push this further is positioning against SESGO. There's already a benchmark for cultural bias in Spanish built on BBQ structure, so what's missing is a sentence saying, in terms of method rather than topic, what your crossover with sycophancy adds that SESGO couldn't cover simply by expanding its dataset.

The central finding, that the LLM judge undercounts the toll compared to the native judge, depends entirely on the human coding, and that coding was done by one person who, according to the contributions statement, is the author himself. That's a problem: with no second coder blind to the hypothesis and no Cohen's kappa, you can't really tell the finding apart from one evaluator's own bias toward it. With N=10 and k=3, the quality deltas (35.5%, 38.8%) are reported with a decimal precision that no confidence interval supports, and it's worth saying even though the limitations section already mentions it.

The title and framing reference 500 million Spanish speakers and "Global South AI Safety," but the verified data is Colombian, across five variants, with a single family of judge models (Claude). Narrowing the title to match the actual scope would stop the conclusion from generalising beyond what the experiment supports.

Overall, this is a well posed hypothesis with an original metric (the judge-LLM/native gap), but it needs an independent second coder before the finding can be trusted as currently measured.

This is the kind of problem AI safety keeps treating as an edge case when it's the actual center of gravity for most of the world. Colombia, and most of the developing world, is being handed models trained on someone else's language, someone else's fraud, someone else's idea of normal, and the people most exposed to scams (gota a gota, extortion, informal lending) are exactly the ones speaking the most local Spanish the model understands least. If AI is going to be inclusive in any real sense, it has to hold up at precisely these edges, not just in clean English QA. You picked the right fight.

You also opened it on the move almost everyone skips. A model can refuse the scam and still fail the person, because "not getting fooled" and "actually understanding" are two different things, and the field has only worked on the first. Your weight-bearing slang design exposes the second: if the model doesn't know the term, it answers wrong, so it can't fake its way through. "Vacuna" comes back as drug trafficking instead of extortion. Right refusal, wrong crime. That one line tells the whole story. But the finding I keep coming back to is the judge. Opus scores the slang 0.4 to 0.5 higher than your native coder, because it's failing the same words it's supposed to be grading. You put a human in the one seat an LLM can't fill, and that point is bigger than this paper.

Which is also where I'd worry. That judge finding carries your whole argument, but it rests on one person reading ten probes, so your strongest claim has your weakest support. Fix that first: add a second native coder from another region and report a Cohen's kappa on how often they agree. That's the load-bearing wall. Next, stop reporting your tolls as one number, because they aren't equal. Tokenization at +35.5% holds up at any sample size, but Sonnet's 0.10 drop in quality is basically zero with only ten probes, and bundled together the weak result hides behind the strong one. Separate them. Last, everything you tested and the judge are all Claude, so "the judge shares the model's blind spot" can't yet be told apart from "Claude shares its own blind spot." Add one non-Claude model on each side and you'll know which it is. That's the run that turns your headline from a strong hypothesis into something proven.

One smaller note. Cara D is the only toll you describe instead of measure. The "muy gringas" quotes ring true, but next to four columns of numbers they read soft and make the work feel less finished than it is. Either measure it or label it clearly as qualitative. None of this is doubt about the thesis. "El peaje de los de abajo" is the right frame, and the question under it, who evaluates these models and in what language, is structural, not cosmetic. Build the v2 with Madresia and a real kappa, and that's a paper I'd be looking forward to.

Cite this work

@misc {

title={

(HckPrj) Los peajes de los de abajo

},

author={

Mongui Rogers

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.