Durinn Calibration
Ryan Marinelli, Victor Strandmoe · Team Durinn
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Durinn analyzes Hacktoberfest repositories to build OWASP-aligned datasets and calibrate a ProtectAI prompt-injection model for detecting vulnerable code. Calibration improves the model’s ability to identify unsafe snippets.
Reviews
The creators of the dataset provides an explanation of why the Raw Protected AI fails distinguishing SAFE and POISONED samples. The statistical testing seems robust, since the distributions are showed in the documentation. Calibration shows to improve the recall in a significant amount for rare events data.
Cite this project
@misc{marinelli2025durinn,
title = {{Durinn Calibration}},
author = {Ryan Marinelli and Victor Strandmoe},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/durinn-calibration-3vub}},
url = {https://apartresearch.com/sprints/projects/durinn-calibration-3vub}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …