HAZE: Adversarial Multi-Agent Scrutiny for Vulnerability Detection
ALEX DANIEL NOVELO FUENTES
HAZE (Hallucination-Aware Zero-sum Examination) is an AI safety project that tackles a failure mode in AI-assisted security auditing: single-agent LLM auditors tend to manufacture findings, flagging safe code as vulnerable. HAZE reframes the task from "is there a bug?" to "did this accusation survive refutation?" using a multi-agent debate protocol—two attacker and two defender models argue over a code artifact while an isolated judge rules only on what withstands cross-examination.
Grounded in AI safety principles (debate as scalable oversight and AI-control-style distrust of any single LLM pass), it was measured on a 10-CVE benchmark: a single-agent baseline flagged 80% of patched code as vulnerable, while HAZE cut that false-positive rate to 20% while keeping 80% recall—a 4× reduction in hallucinated findings, at roughly USD 0.02 per verdict. Built and reproducible, it shows how adversarial multi-agent scrutiny can make AI-driven security pipelines more trustworthy.
HAZE addresses an important problem in LLM-based code auditing by reducing false positives through a structured 2v2 prosecution/defense debate with an independent judge. The benchmark-based evaluation on labeled CVEs, along with comparisons against a single-agent baseline, provides a solid foundation. The reduction in false positives (80% to 20% at 80% recall), thoughtful error analysis, statistical reporting, and transparent discussion of limitations make this a credible and well-executed study. I also appreciated the open-source evaluation harness and the authors' restraint in not overstating statistically underpowered results.
The primary limitation is the small evaluation set, which makes the observed improvement promising but not yet conclusive. Expanding the benchmark would significantly strengthen the statistical claims. The system would also benefit from multi-file context or retrieval support to address cross-file false negatives, and it would be useful to better characterize the vulnerabilities missed when recall drops from 100% to 80%, since false negatives can have a high security impact.
Overall, this is a thoughtful, well-measured paper with a practical approach to improving LLM vulnerability detection, and it provides a strong foundation for future work.
This is the most rigorous submission of the set. The framing — single-agent auditors over-flag ("100% recall, 80% FPR") so reframe the question as "did the accusation survive refutation?" — is sharp, and the work is unusually honest about its own limits. Main issues:
The headline result is underpowered and the team knows it. McNemar p=0.34 at n=20 means the 80%→20% FPR drop is directional, not significant. They state this plainly (good), but the abstract and README still lead with "4× reduction" as a headline. The substance is fine; the framing slightly oversells what the statistics support.
Two different framings/runs are blended. Phase 1 (n=20, VULNERABLE/SEGURO) is the primary, but the error taxonomy and round-ablation come from the web run (n=120, "malware" framing), which the paper itself marks "superseded." Conclusions drawn from a superseded framing shouldn't carry into the headline narrative without a caveat. Re-derive the taxonomy on Phase 1 data.
Figure inconsistencies. Figure 1/Figure 2 show FPR=30 for HAZE, but Table 1 and the abstract report 20%. The figures appear to plot the web run while the text reports Phase 1 — a reader can't tell which number is canonical. Reconcile.
n=20 with the baseline as the comparator is thin. The baseline (100% recall via flagging nearly everything) is almost a strawman — a single agent prompted toward calibration, not just detection, would be a fairer control. Worth noting the baseline is weak by construction.
Judge is Claude Sonnet 4.5 while one debater is also a strong model — possible model-identity effects on judging aren't probed. A judge-swap ablation would strengthen the isolation claim.
Real world scale application was missing how to frame the HAZE with multiple LLM's
I like this one because it goes after a real and annoying problem: ask one model to find a bug and it will invent one to please you. Turning the judge's question into "did the accusation survive a rebuttal" is a clever fix. And you actually measured it on labeled CVEs instead of just shipping a demo, with different models in each seat and a separate judge, which matters. You were straight about the stats too, the p value is not significant and you said so. Two things hold it back. With only 20 cases it is underpowered, so getting to the 60-plus you mention would tell you if the false-positive drop is real. And the single-agent baseline is kind of a punching bag, it over-flags by design, so beating it is not saying much yet. Keep an eye on the recall you are giving up. Good, honest build.
Cite this work
@misc {
title={
(HckPrj) HAZE: Adversarial Multi-Agent Scrutiny for Vulnerability Detection
},
author={
ALEX DANIEL NOVELO FUENTES
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


