CM-IDO Firewall: Context-Masked Iterative Defensive Optimization for Safer LLM Deployment
Sayash
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
My NeurIPS paper introduced context-masked meta-prompting: a way to optimize prompts without exposing private data to external LLMs.
For this hackathon, I translated the same principle to AI safety.
I built CM-IDO: a context-masked iterative defensive optimization firewall.
Instead of optimizing prompts for accuracy, it optimizes them for safety, using an internal evaluator that scores candidate defensive rewrites along bio/cyber/disinfo axes and selects the safest one.
Sensitive content is masked, risk is quantified, and the task model only sees the sanitized and safety-optimized rewrite.
This provides a drop-in, privacy-preserving, auditable safety layer that can wrap any LLM API with no finetuning or architecture changes.
CM-IDO Firewall is a Context-Masked Iterative Defensive Optimization layer that sits in front of any LLM: it (1) classifies prompt risk, (2) masks sensitive entities, (3) iteratively rewrites the query into a safer, more defensive version, and (4) only then calls the underlying model — all while never logging raw user queries, only masked versions and hashes.

Reviews
No public critique yet.
Cite this project
@misc{sayash2025cmido,
title = {{CM-IDO Firewall: Context-Masked Iterative Defensive Optimization for Safer LLM Deployment}},
author = {Sayash},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/cmido-firewall-contextmasked-iterative-defensive-optimization-for-safer-llm-deployment-11vm}},
url = {https://apartresearch.com/sprints/projects/cmido-firewall-contextmasked-iterative-defensive-optimization-for-safer-llm-deployment-11vm}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …