Savaguarda

Tomás Mandl

Salvaguarda is an AI governance toolkit for Global South small and medium-sized enterprises (SMEs), built as a Latin America beta. Because ISO/IEC 42001 and the NIST AI RMF demand expertise and budget most Latin American SMEs lack, the proposal crosswalks down to twelve zero-to-low-cost controls, adds a three-tier risk rubric that scales them to a firm's exposure, and delivers a thirty-minute, privacy-preserving self-assessment that returns a tier verdict and a prioritized, costed action plan with regulatory flags for Brazil, Colombia, and Paraguay.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

The project addresses an important and underexplored AI safety challenge: helping SMEs in the Global South adopt practical AI governance despite limited resources. The motivation is compelling, and the proposed framework successfully translates comprehensive governance standards into a lightweight, actionable approach. The focus on localization, low-resource languages, and regional regulatory contexts is particularly valuable and distinguishes the work from generic governance checklists.

What limits the current submission is that it primarily presents a well-reasoned framework rather than demonstrating that the framework works in practice. The evaluation relies on representative case studies and face-validity arguments, but there is little evidence that SMEs would complete the assessment, implement the recommended controls, or experience measurable improvements in governance or safety outcomes. The next stage of the project would benefit from pilot deployments with real organizations and observations of how the framework changes organizational practices over time.

Several questions also remain about the design choices. Why are these twelve controls the appropriate minimum governance baseline? Are they derived systematically from existing frameworks, or could other subsets achieve similar outcomes? A clearer justification for why these controls—and not others—would strengthen confidence in the methodology.

Similarly, while the three-tier risk model is intuitive, it would be useful to understand how robust it is across a wider variety of AI use cases. For example, are there deployments that fall between tiers or require different governance interventions? Additional validation with domain experts or comparisons against existing governance assessment methods would help establish that the proposed rubric generalizes beyond the illustrative examples presented.

The localization component is one of the strongest aspects of the submission, but it currently remains largely conceptual. Demonstrating that localized language testing uncovers risks that existing governance guidance would otherwise miss would make the contribution considerably stronger and better establish its novelty.

One possible direction for future work would be to structure the evaluation around one or two detailed organizational case studies. Showing how an SME initially assessed its AI use, implemented the recommended controls, encountered practical challenges, and changed its governance processes over time would provide stronger evidence of the framework's real-world value than hypothetical walkthroughs alone.

Overall, this is a thoughtful and well-presented contribution that identifies a genuine gap in AI governance. With empirical validation, a working implementation, and evidence of organizational adoption, it has the potential to become a valuable practical resource for AI safety in resource-constrained environments.

3/2/4

Criteria 1 - Impact Potential & Innovation: 3.5

Criteria 2 - Execution Quality: 2

Criteria 3 - Presentation & Clarity: 4

Identifies a genuinely under-attended problem , well-structured, cleanly enumerated threat model, useful tables, explicit tier logic with an appendix, and a thorough, helpful dual-use/limitations section. The crux of the problem is that the artifact does not exist.

A sharp and well-framed project. The "long tail of small deployers" framing is compelling, and the writing is excellent. The only gap I see, which I believe can be quickly fixed, is that the actual tool has not beein built or tested. Building a reference implementation and piloting it with a couple of real SME's will substantially strengthen the work!

Cite this work

@misc {

title={

(HckPrj) Salvaguarda

},

author={

Tomás Mandl

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.