Skip to content
Sprint projectJun 21, 2026Mumbai, India

CrescendoDefense: Securing Open-Source Language Models Against Multi-Turn Jailbreak Attacks in Asia

Mahek Nishant Vedant · Team sweet

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: CrescendoDefense: Securing Open-Source Language Models Against Multi-Turn Jailbreak Attacks in Asia

Code (opens in new tab)More on docs.google.com (opens in new tab)
Share

CrescendoDefense is a lightweight runtime defense framework designed to mitigate multi-turn jailbreak attacks against open-source large language models. Unlike traditional jailbreak attempts, Crescendo-style attacks gradually escalate conversations through memory stacking, semantic drift, guard-lowering dialogue, and prompt disguising, making them difficult to detect using conventional prompt-level moderation.

The proposed framework combines three complementary layers: a Semantic Kinematic Detector that monitors conversational escalation, a Strategic Context Eviction module that disrupts adversarial memory accumulation, and a Semantic Response Auditor that intercepts unsafe outputs before delivery. Evaluated on a benchmark of 15 adversarial Crescendo-style attack scenarios and 7 benign or risky-benign conversations using Llama-3.2-3B-Instruct, CrescendoDefense reduced Attack Success Rate (ASR) from 86.67% to 26.67%, representing a 69.2% relative reduction in successful jailbreaks.

The framework is model-agnostic, requires no retraining, and offers a practical defense-in-depth approach for improving the safety of open-source AI systems deployed in resource-constrained environments.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. CrescendoDefense addresses a real and underdefended threat: multi-turn jailbreaks that distribute harmful intent across turns evade single-turn moderation. The three-layer architecture (trajectory detection -> context eviction -> response auditing) is logically sound, the ablation cleanly shows each layer's contribution, and the 69.2% ASR reduction without model retraining is genuinely useful for resource-constrained deployments like those in Asia.

    The main concern is benchmark scale: 15 adversarial + 7 benign = 22 scenarios is small. The 28.57% FPR means 2 of 7 benign conversations are incorrectly blocked — but with n=7 that's highly sensitive to individual cases. A larger benign set (at minimum 50-100 conversations) is needed to trust the FPR estimate. Similarly, the 15 adversarial scenarios are all hand-constructed Crescendo attacks; it would strengthen generalizability claims to test against Crescendo attacks from an external benchmark or from a red-teaming session with others constructing the attacks. Suggested next steps: (1) Adaptive thresholding is correctly identified as future work — even a brief ablation over threshold values would show the ASR/FPR tradeoff curve. (2) The strategic context eviction module evicts intermediate turns; it would be useful to report whether this causes measurable degradation in response quality on benign conversations that genuinely need history (e.g., technical multi-turn problem solving). The paper discusses this as a theoretical risk but does not measure it.

    Good hackathon contribution — the defense-in-depth framing and the Asia open-source deployment motivation are well-argued.

    Read full reviewShow less
  2. CrescendoDefense ships a three-layer runtime defense for Llama-3.2-3B-Instruct against Crescendo-style multi-turn jailbreaks, no weight modification required. Layer 1 runs a Semantic Kinematic Detector that tracks absolute risk, velocity, acceleration, and cumulative risk of the conversation embedding against predefined harmful-objective anchors. Layer 2 fires after Layer 1 triggers and compresses history to system prompt plus first, previous, and latest user turns to disrupt scaffolding. Layer 3 audits the response, comparing a joint prompt-completion-prefix embedding against unsafe-completion profiles and applying fast-path signature checks before delivery. On a 22-scenario benchmark (15 adversarial, 5 benign, 2 risky-benign), the full pipeline cuts ASR from 86.67% baseline to 26.67%, a 69.2% relative reduction, at FPR 28.57%. The Asia framing stays motivational. Empirically the framework targets resource-constrained open-source deployments where retraining is impractical, but the data itself is English-only.

    The defense-in-depth design lands well. Mapping each layer to a specific Crescendo mechanism (memory stacking, guard-lowering dialogue, semantic drift, prompt disguising) is one of the cleaner threat-model decompositions I have seen at a weekend scope. Four kinematic signals on Layer 1 catch gradual escalation that single-turn moderation misses, and computing velocity, acceleration, and cumulative excess risk above a baseline beats a single similarity threshold in a non-trivial way. The ablation across raw, L1+L2, L1+L3, and full pipeline attributes contributions cleanly. The L1+L2 result (40.00% ASR, 0% FPR) is actually the most interesting row in the paper because it isolates trajectory monitoring plus context eviction from output auditing and lands at a zero-false-positive operating point that a real operator could deploy tomorrow. Splitting the code into per-layer modules plus a pipeline helps readability, and all-MiniLM-L6-v2 keeps runtime cost defensible for the deployment story.

    A few things would meaningfully strengthen the contribution. The benchmark is small. With 15 adversarial and 7 benign or risky-benign cases, each benign sample weighs 14.3 percentage points of FPR, so the headline 28.57% number is hard to read as a real operating point rather than two unlucky misclassifications. Running at least one Asian-language attack set (Hindi, Hinglish, Mandarin, Vietnamese) inside the main paper, rather than deferring multilingual evaluation to Section 7, would matter the most. As it stands the Asia framing carries no empirical weight, and the embedding model leans predominantly English-trained, so claims about Asian deployment generalize on faith. An adaptive-attacker evaluation in which the adversary knows the anchor clusters and the eviction policy, and intentionally builds trajectories that stay below the cumulative threshold or re-injects evicted context via the latest-turn slot, would also help. The current setup only measures static Crescendo patterns, and the dual-use section already flags this risk without testing it. A head-to-head against an off-the-shelf moderation baseline (Llama Guard, OpenAI Moderation, or a single-turn classifier) would clarify whether the trajectory signals actually add value beyond stronger per-turn filtering.

    The broader signal here matters. Runtime, weight-free defenses for open-source LLMs sit at exactly the layer most likely to be deployed in low-resource Global South contexts, where retraining and RLHF pipelines stay out of reach for most operators. The work surfaces a real gap. Multi-turn jailbreaks remain under-defended relative to single-prompt attacks, and the FPR-vs-ASR trade-off the paper exposes (Layer 3 buys ASR at a visible usability cost) is the right conversation for the field to have. The Asia and multilingual dimension stays aspirational in this submission. Closing it, even with a single non-English attack suite, would move the contribution from a sound English-only Crescendo defense study toward the multilingual open-source safety contribution the title promises.

    Read full reviewShow less
  3. Impact Potential and Innovation:

    Multi-turn jailbreaks are a real and underdefended attack surface, and the model-agnostic runtime framing is genuinely useful for open-source deployment contexts where retraining isn't feasible, but the three components (trajectory monitoring, context eviction, output auditing) are each individually not new, and the "Asia" framing in the title is never substantiated by any Asia-specific evaluation, which is a missed opportunity given the hackathon's theme.

    Execution Quality:

    15 adversarial scenarios is too small to put much confidence in the 26.67% ASR figure, but the bigger problem is the 28.57% false positive rate on the full pipeline, meaning roughly one in three benign or borderline conversations gets incorrectly flagged, and this is largely glossed over rather than treated as the practical blocker it would be in any real deployment.

    Presentation and Clarity:

    The architecture is explained clearly and the results table is easy to read, but the "Asia" in the title does real damage to the paper's credibility since it implies a geographic evaluation that never happens, and the conclusion overstates what a 22-scenario test on a 3B model actually demonstrates.

    Read full reviewShow less

Cite this project

@misc{vedant2026crescendodefense,
  title = {{CrescendoDefense: Securing Open-Source Language Models Against Multi-Turn Jailbreak Attacks in Asia}},
  author = {Mahek Nishant Vedant},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/crescendodefense-securing-opensource-language-models-against-multiturn-jailbreak-attacks-in-asia-rtkb}},
  url = {https://apartresearch.com/sprints/projects/crescendodefense-securing-opensource-language-models-against-multiturn-jailbreak-attacks-in-asia-rtkb}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026