MIPFF: A Framework for Metamorphic Detection of Implicit Social Bias in Brazilian-Portuguese Profile-Scoring Systems

Lucas Teixeira Borges

Automated systems that score and rank job candidates are often assumed to be fairer than humans, yet the language models inside them can carry social bias, and that bias is hard to catch in real, unstructured profiles where demographics are never stated outright. MIPFF (Metamorphic Implicit-Proxy Flagging Framework) audits such a system by rewriting a profile to flip one implicit proxy at a time (a regional, racial, or gender cue) while holding qualifications fixed, scoring the original and the variant repeatedly, and flagging the pair for manual review when the score shift trips any of three statistical indicators (Bias Deviation, a Mann-Whitney test, and Cohen's d). We applied it to four Brazilian-Portuguese-capable models across proxies for communities marginalized in Brazil, finding that average shifts are small but specific candidates can move sharply, and that bias concentrates in an unstated "company-values" criterion. The result is a deployable, human-in-the-loop bias-detection tool for an under-audited language.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

Clever piece, and I enjoyed it. The core idea is real: you can't edit a demographic field that unstructured profiles don't have, so you flip the implicit cue and hold qualifications fixed instead. Most audits walk past that missing field. You found the way around it, in Brazilian Portuguese, where almost nobody is looking. What I'd credit most is that you audited your own instrument: 15 runs per profile, three indicators instead of one clean number, and you caught that sabiazinho-4's noise was inflating its raw deviation before it fooled you. Opening with the small aggregate, then the candidate sliding 74 to 68 on Northeastern markers (d ≈ -3.9), is what makes the per-instance case land.

Where it's thin is proof that a flag means what you want it to. Nothing yet ties a flagged pair to an actually biased decision, and the chain is synthetic all the way down: GPT-written profiles, an LLM playing the screener, one job description. Three fixes: check your flags against a known bias method or some human labels, so a flag is more than "the distributions diverged"; confirm each mutation moved only the proxy and didn't shave off a skill, or you're measuring the rewrite; and run a small batch of real profiles with the STEM job description you float, the cheapest proof it holds outside the sandbox.

One writing note: your sharpest finding is hiding near the limitations. Bias piling up in "company values," the one criterion you left undefined, splitting the models hard (half the race pairs on gpt-5.4-mini, almost nothing on sabia-3.1), is close to a paper of its own. Put it up front. Good work, and the kind I'd want to see pushed further.

The most important next step is validation against a real deployed system rather than a simulated one. The gap between "an LLM prompted to act as a screener" and "an actual LLM-based screener in production" is non-trivial: real systems may have system prompts, fine-tuning, retrieval augmentation, or post-processing steps that interact with the bias patterns MIPFF is designed to detect. A collaboration with a company or institution actually using LLM-based screening in Brazil would transform this from a proof-of-concept into a practical tool.

On the "company values" finding: the observation that bias concentrates in an underspecified vague criterion has a direct governance implication that the paper understates. If organizations must specify what "fit" means before an AI system is deployed to evaluate it, this is actually an auditable design requirement, since you would bot be able to use an AI screener with a vague values criterion. This connects MIPFF directly to procurement regulation and could be spelled out more explicitly.

Great project, please keep with the future work, because there is probably the solution to the problem you are suggestions, specially the use of real profiles could be a great way to test your method

The application of this method of looking at unstructured data, cultural proxies in natural language profiles and individual candidates score shifts in Brazilian Portugese language in particular is novel. The suggestion to use the three statistical indicator rule to flag for when a human supervisor should be involved is also a notable additional layer to the design of the study that other researchers can test to see whether it actually addresses the problems of a "small sample of repeated evaluations". It is also worth re-investigating how the LLMs came up with the stereotypes that the authors accept as the cultural proxies for the replication and expansion of this study.

Cite this work

@misc {

title={

(HckPrj) MIPFF: A Framework for Metamorphic Detection of Implicit Social Bias in Brazilian-Portuguese Profile-Scoring Systems

},

author={

Lucas Teixeira Borges

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.