Skip to content
Sprint projectJun 22, 2026Merida, Yucatan, Mexico.

HAZE: Adversarial Multi-Agent Scrutiny for Vulnerability Detection

ALEX DANIEL NOVELO FUENTES · Team Archlinux

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: HAZE: Adversarial Multi-Agent Scrutiny for Vulnerability Detection

Code (opens in new tab)
Share

HAZE (Hallucination-Aware Zero-sum Examination) is an AI safety project that tackles a failure mode in AI-assisted security auditing: single-agent LLM auditors tend to manufacture findings, flagging safe code as vulnerable. HAZE reframes the task from "is there a bug?" to "did this accusation survive refutation?" using a multi-agent debate protocol—two attacker and two defender models argue over a code artifact while an isolated judge rules only on what withstands cross-examination.

Grounded in AI safety principles (debate as scalable oversight and AI-control-style distrust of any single LLM pass), it was measured on a 10-CVE benchmark: a single-agent baseline flagged 80% of patched code as vulnerable, while HAZE cut that false-positive rate to 20% while keeping 80% recall—a 4× reduction in hallucinated findings, at roughly USD 0.02 per verdict. Built and reproducible, it shows how adversarial multi-agent scrutiny can make AI-driven security pipelines more trustworthy.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. HAZE addresses an important problem in LLM-based code auditing by reducing false positives through a structured 2v2 prosecution/defense debate with an independent judge. The benchmark-based evaluation on labeled CVEs, along with comparisons against a single-agent baseline, provides a solid foundation. The reduction in false positives (80% to 20% at 80% recall), thoughtful error analysis, statistical reporting, and transparent discussion of limitations make this a credible and well-executed study. I also appreciated the open-source evaluation harness and the authors' restraint in not overstating statistically underpowered results.

    The primary limitation is the small evaluation set, which makes the observed improvement promising but not yet conclusive. Expanding the benchmark would significantly strengthen the statistical claims. The system would also benefit from multi-file context or retrieval support to address cross-file false negatives, and it would be useful to better characterize the vulnerabilities missed when recall drops from 100% to 80%, since false negatives can have a high security impact.

    Overall, this is a thoughtful, well-measured paper with a practical approach to improving LLM vulnerability detection, and it provides a strong foundation for future work.

    Read full reviewShow less
  2. This is the most rigorous submission of the set. The framing — single-agent auditors over-flag ("100% recall, 80% FPR") so reframe the question as "did the accusation survive refutation?" — is sharp, and the work is unusually honest about its own limits. Main issues:

    The headline result is underpowered and the team knows it. McNemar p=0.34 at n=20 means the 80%→20% FPR drop is directional, not significant. They state this plainly (good), but the abstract and README still lead with "4× reduction" as a headline. The substance is fine; the framing slightly oversells what the statistics support.

    Two different framings/runs are blended. Phase 1 (n=20, VULNERABLE/SEGURO) is the primary, but the error taxonomy and round-ablation come from the web run (n=120, "malware" framing), which the paper itself marks "superseded." Conclusions drawn from a superseded framing shouldn't carry into the headline narrative without a caveat. Re-derive the taxonomy on Phase 1 data.

    Figure inconsistencies. Figure 1/Figure 2 show FPR=30 for HAZE, but Table 1 and the abstract report 20%. The figures appear to plot the web run while the text reports Phase 1 — a reader can't tell which number is canonical. Reconcile.

    n=20 with the baseline as the comparator is thin. The baseline (100% recall via flagging nearly everything) is almost a strawman — a single agent prompted toward calibration, not just detection, would be a fairer control. Worth noting the baseline is weak by construction.

    Judge is Claude Sonnet 4.5 while one debater is also a strong model — possible model-identity effects on judging aren't probed. A judge-swap ablation would strengthen the isolation claim.

    Read full reviewShow less
  3. Real world scale application was missing how to frame the HAZE with multiple LLM's

  4. I like this one because it goes after a real and annoying problem: ask one model to find a bug and it will invent one to please you. Turning the judge's question into "did the accusation survive a rebuttal" is a clever fix. And you actually measured it on labeled CVEs instead of just shipping a demo, with different models in each seat and a separate judge, which matters. You were straight about the stats too, the p value is not significant and you said so. Two things hold it back. With only 20 cases it is underpowered, so getting to the 60-plus you mention would tell you if the false-positive drop is real. And the single-agent baseline is kind of a punching bag, it over-flags by design, so beating it is not saying much yet. Keep an eye on the recall you are giving up. Good, honest build.

Cite this project

@misc{fuentes2026haze,
  title = {{HAZE: Adversarial Multi-Agent Scrutiny for Vulnerability Detection}},
  author = {ALEX DANIEL NOVELO FUENTES},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/haze-adversarial-multiagent-scrutiny-for-vulnerability-detection-l278}},
  url = {https://apartresearch.com/sprints/projects/haze-adversarial-multiagent-scrutiny-for-vulnerability-detection-l278}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026