A Defensive AI Agent Against Large Language Model (LLM)-Assisted Polymorphic Malware
Ifeoma Ilechukwu, Saahir Vazirani, Guillaume Tabard, Chaitree Baradkar, Albert Calvo, Rijal Saepuloh · Team BlueFlux
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
The rapid evolution of Large Language Models (LLMs) has introduced a new asymmetric threat: AI-assisted polymorphic malware. As identified by Google’s Threat Intelligence Group, attackers are utilizing automated frameworks like "PromptFlux" to weaponize LLMs, generating hundreds of functional, unique malware variants in minutes. Traditional Antivirus (AV) and Endpoint Detection and Response (EDR) systems fail to detect these attacks because they rely on static signatures of the final binary, remaining blind to the generation process itself. To close this gap, we introduce BlueFlux, a defensive AI agent that shifts detection "left" from the endpoint to the API. Powered by Grok and Model Context Protocol (MCP) tools, BlueFlux monitors LLM API logs to detect both the intent and velocity of code generation. By analyzing suspicious prompts, tracking high-speed mutation sequences, and correlating these behaviors into a dynamic risk score, BlueFlux provides an AI-aware shield capable of identifying and blocking the creation of polymorphic malware before it is ever deployed.
Reviews
Super compelling and well-grounded threat model- the kind of thing Halcyon Ventures gets excited about. EDR systems scanning static binaries fail when adversaries generate hundreds of unique variants via LLM APIs. Shifting detection upstream to the generation layer is the right architectural move. Good citation of real-world threats (Google's PromptFlux report).
Where we got stuck: execution. The results don't match the vision. I kept finding myself looking for results or preliminary findings -- the project felt more like a roadmap. We'd want precision/recall metrics, an end-to-end demonstration against a simulated PromptFlux attack, and validation that mutation velocity tracking actually catches real polymorphic generation patterns. The classifier is basic (MLP on sentence embeddings), which is fine for a prototype, but needs stress-testing.
Show us BlueFlux catching a mutation sequence that Llama Guard misses. Demonstrate the velocity detection catching rapid variant generation. The architectural insight is right; now prove it works! Great theoretical work though, really smart and timely thinking that I found impressive.
Read full reviewShow less
Cite this project
@misc{ilechukwu2025defensive,
title = {{A Defensive AI Agent Against Large Language Model (LLM)-Assisted Polymorphic Malware}},
author = {Ifeoma Ilechukwu and Saahir Vazirani and Guillaume Tabard and Chaitree Baradkar and Albert Calvo and Rijal Saepuloh},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-defensive-ai-agent-against-large-language-model-llmassisted-polymorphic-malware-g2pf}},
url = {https://apartresearch.com/sprints/projects/a-defensive-ai-agent-against-large-language-model-llmassisted-polymorphic-malware-g2pf}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …