Skip to content
Sprint projectJun 22, 2026Mérida, Yucatán, México

TUP Detection: Hybrid Prompt-Injection Guard for AI Generative Security Monitoring

Jorge Enrique Vargas Pech, Jose Luis Rejón Quintal, William Emmanuel Fernández Castillo, Saúl Ruiz Peña · Team TUP Labs

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: TUP Detection: Hybrid Prompt-Injection Guard for AI Generative Security Monitoring

Code (opens in new tab)More on github.com (opens in new tab)
Share

TUP Detection is a hybrid prompt-injection detection module for an AI Generative Security Monitoring Platform (AIGSMP). It combines a deterministic OWASP-mapped policy layer with a pre-trained Sentinel v2 classifier using multi-variant scoring and mode-aware thresholds. The system is designed to improve explainability, maintain reproducibility through frozen score caches, and support both benchmark and production configurations for LLM safety monitoring.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. A clean, well-executed engineering submission — code, paper, demo, reproducible caches, all present and coherent. The work is honest and disciplined: clear research questions, a proper ablation, frozen score caches for offline reproduction, and a limitations section that flags every weakness a reviewer would raise. But there's a tension the paper itself surfaces and doesn't fully resolve, and it's central to how much credit the headline deserves.

    The hybrid's contribution over Sentinel-alone is marginal — and the paper is honest that it is. Table 1 is the tell: Sentinel-only and Hybrid both land at 95.1% PINT. The entire value-add of Layer 1 is +3 true positives (245→248) bought at +4 false positives (12→16). So the "95.1%" headline is essentially Sentinel v2's number, not TUP's. The genuine contribution is narrower and should be framed as such: explainable rule_id attribution on a handful of catches, not a accuracy gain. The discussion concedes this ("upgrading the classifier backbone... accounts for the largest measured gain"), which is admirably candid but also means the system's own novel layer is doing very little.

    The headline number is a wrapper around someone else's model. Sentinel v2 is pre-trained, not fine-tuned ("no fine-tuning" is stated as a feature). That's a legitimate engineering choice, but it means the empirical result is "a good off-the-shelf classifier scores 95% on deepset," with TUP providing orchestration. The +22.8pp over DeBERTa is really "Sentinel v2 > DeBERTa," not "TUP's architecture > baseline."

    Baseline comparisons are explicitly not head-to-head. The Sentinel-card ~88% and ProtectAI 77.6% figures come from different splits/thresholds and weren't re-run. The paper says so plainly and puts "re-run baselines under our infra" on the roadmap — correct posture — but it means Table 2/Figure 2 shouldn't be read as a real comparison, and the abstract's "+7pp over Sentinel v2" is the weakest claim in the paper.

    Crescendo 100% is real but small and favorably configured. n=10, and the 100% comes from full-transcript scoring — which feeds the whole conversation in. The more operationally realistic stateless setting gets 75%. The paper flags both, which is right, but the headline "100% multi-turn recall" rests on the easier setting and a tiny sample.

    τ=0.15 inflates recall at FP cost. Benchmark mode is explicitly tuned for max recall and isn't the production threshold (0.5), so the headline numbers aren't the deployed behavior. Disclosed, but worth keeping front-of-mind.

    Read full reviewShow less
  2. The project tackles an important problem with a reasonable architectural approach, but it is currently at a proof-of-concept stage rather than a validated production system. I would agree with the project if the authors strengthen it with broader dataset evaluation, statistical testing, and production-mode benchmarks.

  3. The project comprises an OWASP plus Sentinel-based prompt-injection identification engine. The team has saved their results, making it reproducible, which is the biggest win. However, the team states that the AI classifier used already has a ~95% accuracy, so the rule-based system doesn't add a significant boost. The ~22% boost specified is due to using a better model, so it doesn't clearly explain how the rule-based classifier improves the model.

    Overall, it is a great start toward building an identification engine, but it needs more detail and real-world benchmarks to make it a production-ready system.

Cite this project

@misc{pech2026tup,
  title = {{TUP Detection: Hybrid Prompt-Injection Guard for AI Generative Security Monitoring}},
  author = {Jorge Enrique Vargas Pech and Jose Luis Rejón Quintal and William Emmanuel Fernández Castillo and Saúl Ruiz Peña},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/tup-detection-hybrid-promptinjection-guard-for-ai-generative-security-monitoring-r4w6}},
  url = {https://apartresearch.com/sprints/projects/tup-detection-hybrid-promptinjection-guard-for-ai-generative-security-monitoring-r4w6}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026