Contestabilidad algorítmica en el Estado colombiano: un canal de objeción asistido por IA

Emely Condor, Federico Perez

Los sistemas algorítmicos del Estado colombiano ya toman o asisten decisiones que afectan derechos, pero los canales para objetar esas decisiones casi no existen. Construimos un Índice de Contestabilidad Ciudadana que operacionaliza, artículo por artículo, la Directiva Conjunta 007 de 2025 de la Procuraduría y la Defensoría, y lo aplicamos a 409 sistemas de decisión automatizada del repositorio de la Universidad de los Andes; sistemas de alto impacto se verificaron contra las webs oficiales de cada entidad. El resultado central es una brecha medible: la transparencia informacional es alta (las entidades publican qué es el sistema, su objetivo y sus datos) mientras la contestabilidad es casi nula. Solo 6 de 409 sistemas ofrecen un canal de objeción completo, 391 no ofrecen ninguno, y ninguno de los 154 sistemas de alto riesgo publica un análisis de impacto algorítmico. A partir de esta evidencia proponemos: el diseño de un protocolo de canal de objeción asistido por IA que recibe el reclamo de un ciudadano, lo clasifica por tipo de sistema y lo enruta al responsable según la Directiva 007. La conclusión para la seguridad de la IA es que la contestabilidad debe diseñarse antes de que los agentes IA lleguen, no después.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is a clear and policy-relevant project. Its most valuable contribution is the Contestability Index applied to 409 automated decision systems in the Colombian public sector. The finding is concrete and actionable. Public entities often disclose what systems exist and what they do, but almost never provide a meaningful channel for affected citizens to object. The contrast between informational transparency and practical contestability is the paper’s strongest insight, and the figures and tables communicate it effectively.

The proposed AI-assisted objection channel is useful as a governance design pattern. The routing architecture (receiving a citizen complaint, identifying the relevant system, classifying it under risk categories, routing it to the correct institutional contact, and preserving traceability) is practical and well aligned with the paper’s diagnosis. The Appendix A proposal that contestability should “follow the function” when an AI agent replaces a human official is promising and could be developed into a model contractual clause for public procurement.

The first limitation is methodological. For example, the core empirical claim depends heavily on whether missing published information is correctly interpreted as lack of contestability. It should be separated from the stronger claim that no objection channel exists in practice.

The second limitation is the connection to AI safety. The paper’s immediate contribution is best understood as algorithmic accountability and public-sector redress. Its AI safety relevance comes from the future scenario in which autonomous agents replace human officials and make contestability structurally harder. That argument is plausible, but underdeveloped. The paper would be stronger if it specified what changes when the decision-maker is an AI agent rather than a conventional automated system. Right now, this is a compelling policy intuition rather than a fully developed safety argument.

For future work, the most important next step is to build and test the objection-channel prototype with realistic or simulated citizen complaints, measuring relevant outcomes.

the paper addresses a critical safety problem, bravo. Methodology is also impressively thorough for a hackathon project, and it's easy to follow. The gap is that this pilot hasn't been used by "real users" limiting it's applicability (and very understandable given the timeframe) but it does slightly limit it's impact, still amazing work.

Este es un proyecto muy sólido, con una contribución clara y directamente relevante para la gobernanza y seguridad de la IA en el sector público: medir la brecha entre transparencia algorítmica y contestabilidad efectiva. La construcción del Índice de Contestabilidad Ciudadana y su aplicación a 409 sistemas del Estado colombiano muestran una ejecución fuerte para un hackathon, especialmente porque el análisis combina codificación sistemática, verificación de casos de alto impacto y una propuesta práctica de canal de objeción asistido por IA. La conexión con agentes de IA es pertinente y valiosa, aunque en algunos momentos podría desarrollarse con mayor precisión técnica para distinguir mejor entre sistemas automatizados actuales y futuros agentes autónomos. La presentación es clara y convincente, con resultados fáciles de interpretar; para fortalecerlo aún más, sería útil profundizar en cómo se implementaría y evaluaría el prototipo del canal de objeción en un piloto real con una entidad pública.

Cite this work

@misc {

title={

(HckPrj) Contestabilidad algorítmica en el Estado colombiano: un canal de objeción asistido por IA

},

author={

Emely Condor, Federico Perez

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.