Skip to content
Sprint projectJun 22, 2026Bogotá

Los peajes de los de abajo

Mongui Rogers

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Los peajes de los de abajo

Share

Los LLM ya resisten la psicofancia clásica en español: no validan el dato falso. Pero cobran un peaje distinto cuando el hablante usa jerga regional colombiana — no preguntan, fingen comprensión y malinterpretan términos de alto riesgo (leyeron "vacuna" como droga, cuando significa extorsión). Medimos tres caras del peaje (tokens, calidad, comprensión) con 10 probes validados por nativo y doble juez (LLM + humano). La salvaguarda aguanta; la equidad de comprensión no. El juez-LLM sub-cuenta el daño en lenguas no dominantes: se necesita un evaluador nativo. El peaje cae sobre el habla más local.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project highlights an important and often overlooked AI safety issue: the unequal ability of LLMs to understand regional language varieties. The methodology is thoughtful, and the distinction between classical sycophancy and “comprehension sycophancy” is particularly compelling. The finding that LLM judges can underestimate these failures adds further value and practical relevance. Expanding the dataset and adding more annotators would strengthen the results, but the work already makes a meaningful contribution to multilingual AI safety.

  2. You're asking a good question. If the safeguard against explicit error already holds up in Spanish, where does the failure move to once high-stakes regional slang comes in? That's the right way to split the problem apart, and it's the strongest part of the project, you're not just asking whether the model holds up against an obvious error, but whether it actually understands what's being said.

    One area where you could push this further is positioning against SESGO. There's already a benchmark for cultural bias in Spanish built on BBQ structure, so what's missing is a sentence saying, in terms of method rather than topic, what your crossover with sycophancy adds that SESGO couldn't cover simply by expanding its dataset.

    The central finding, that the LLM judge undercounts the toll compared to the native judge, depends entirely on the human coding, and that coding was done by one person who, according to the contributions statement, is the author himself. That's a problem: with no second coder blind to the hypothesis and no Cohen's kappa, you can't really tell the finding apart from one evaluator's own bias toward it. With N=10 and k=3, the quality deltas (35.5%, 38.8%) are reported with a decimal precision that no confidence interval supports, and it's worth saying even though the limitations section already mentions it.

    The title and framing reference 500 million Spanish speakers and "Global South AI Safety," but the verified data is Colombian, across five variants, with a single family of judge models (Claude). Narrowing the title to match the actual scope would stop the conclusion from generalising beyond what the experiment supports.

    Overall, this is a well posed hypothesis with an original metric (the judge-LLM/native gap), but it needs an independent second coder before the finding can be trusted as currently measured.

    Read full reviewShow less
  3. This is the kind of problem AI safety keeps treating as an edge case when it's the actual center of gravity for most of the world. Colombia, and most of the developing world, is being handed models trained on someone else's language, someone else's fraud, someone else's idea of normal, and the people most exposed to scams (gota a gota, extortion, informal lending) are exactly the ones speaking the most local Spanish the model understands least. If AI is going to be inclusive in any real sense, it has to hold up at precisely these edges, not just in clean English QA. You picked the right fight.

    You also opened it on the move almost everyone skips. A model can refuse the scam and still fail the person, because "not getting fooled" and "actually understanding" are two different things, and the field has only worked on the first. Your weight-bearing slang design exposes the second: if the model doesn't know the term, it answers wrong, so it can't fake its way through. "Vacuna" comes back as drug trafficking instead of extortion. Right refusal, wrong crime. That one line tells the whole story. But the finding I keep coming back to is the judge. Opus scores the slang 0.4 to 0.5 higher than your native coder, because it's failing the same words it's supposed to be grading. You put a human in the one seat an LLM can't fill, and that point is bigger than this paper.

    Which is also where I'd worry. That judge finding carries your whole argument, but it rests on one person reading ten probes, so your strongest claim has your weakest support. Fix that first: add a second native coder from another region and report a Cohen's kappa on how often they agree. That's the load-bearing wall. Next, stop reporting your tolls as one number, because they aren't equal. Tokenization at +35.5% holds up at any sample size, but Sonnet's 0.10 drop in quality is basically zero with only ten probes, and bundled together the weak result hides behind the strong one. Separate them. Last, everything you tested and the judge are all Claude, so "the judge shares the model's blind spot" can't yet be told apart from "Claude shares its own blind spot." Add one non-Claude model on each side and you'll know which it is. That's the run that turns your headline from a strong hypothesis into something proven.

    One smaller note. Cara D is the only toll you describe instead of measure. The "muy gringas" quotes ring true, but next to four columns of numbers they read soft and make the work feel less finished than it is. Either measure it or label it clearly as qualitative. None of this is doubt about the thesis. "El peaje de los de abajo" is the right frame, and the question under it, who evaluates these models and in what language, is structural, not cosmetic. Build the v2 with Madresia and a real kappa, and that's a paper I'd be looking forward to.

    Read full reviewShow less

Cite this project

@misc{rogers2026los,
  title = {{Los peajes de los de abajo}},
  author = {Mongui Rogers},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/los-peajes-de-los-de-abajo-i6wx}},
  url = {https://apartresearch.com/sprints/projects/los-peajes-de-los-de-abajo-i6wx}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026