Evaluación holística y multi-metodológica de propuestas de seguridad en IA
Mateo Acosta
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
holistic-eval
Evaluación holística y multi-metodológica de propuestas de seguridad en IA
Ante incidentes como el de OpenAI–Hugging Face, el cuello de botella no es la falta de propuestas, sino el tiempo disponible para evaluarlas con rigor desde varios ángulos.
holistic-eval es una herramienta de línea de comandos que evalúa un documento de propuesta frente a una clase de incidente mediante cuatro metodologías independientes ejecutadas en paralelo: prudencia, yellow teaming, red teaming y diseño sinérgico.
Un agregador, orquestado con LangGraph, reconcilia los cuatro informes en un veredicto único. La herramienta corre con Claude y aplica un tope de gasto duro y persistente.
En dos ejecuciones reales, produjo cuatro críticas diferenciadas y un veredicto reconciliado por aproximadamente US$0,40 por evaluación. El hallazgo principal es de viabilidad: resulta barato y reproducible obtener perspectivas metodológicamente distintas de forma automatizada.
Pruébenlo con propuestas del hackathon: github.com/mateo3264/holistic-eval
Uso:
uv sync cp .env.example .env
Agrega tu clave en el archivo .env local:
ANTHROPIC_API_KEY=sk-ant-
Luego ejecuta:
uv run holistic-eval su_propuesta.pdf
La herramienta acepta documentos en PDF, Markdown o texto.
No subas ANTHROPIC_API_KEY al repositorio, al README ni a un chat público. Solicita la clave al responsable del equipo por un canal privado y asegúrate de que .env esté incluido en .gitignore.
Por defecto usa Claude Haiku 4.5, una opción económica de aproximadamente US$0,40 por evaluación.
Para priorizar calidad, puedes cambiar a Sonnet 5 en .env:
LLM_MODEL=claude-sonnet-5 PRICE_PER_MTOK_INPUT=<precio publicado de entrada de Sonnet 5> PRICE_PER_MTOK_OUTPUT=<precio publicado de salida de Sonnet 5>
Es importante configurar PRICE_PER_MTOK_INPUT y PRICE_PER_MTOK_OUTPUT con las tarifas reales de Sonnet 5. El cálculo y cumplimiento del tope de gasto depende de esos valores; conservar las tarifas de Haiku (1/5) subestimaría el gasto.
Reviews
Fairly interesting idea, but without any kind of output quality analysis, it's hard to judge how useful the eval is.
Please next time for an international hackathon, submit the project in English as not everyone speaks Spanish.
Cite this project
@misc{acosta2026evaluacion,
title = {{Evaluación holística y multi-metodológica de propuestas de seguridad en IA}},
author = {Mateo Acosta},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/evaluacin-holstica-y-multimetodolgica-de-propuestas-de-seguridad-en-ia-j9pf}},
url = {https://apartresearch.com/sprints/projects/evaluacin-holstica-y-multimetodolgica-de-propuestas-de-seguridad-en-ia-j9pf}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …