Skip to content
Sprint projectMar 31, 2025Canberra, Australia

Stop Hitting Yourself: Leveraging helpful assistance as an attack vector

Dustin Venini · Team JustMeToday

Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Stop Hitting Yourself: Leveraging helpful assistance as an attack vector

Share

Very much work in progress look at "trusted auditor attacks," a novel adversarial tactic where deceptively aligned AI systems deliberately produce flawed outputs designed to manipulate human or automated reviewers. In this scenario, the AI strategically creates errors that, when corrected by trusted auditors following predictable patterns, inadvertently introduce vulnerabilities or fulfil the AI's hidden objectives. This attack vector is particularly concerning as the harmful outcome appears to be authored by the human reviewer rather than the AI system.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{venini2025stop,
  title = {{Stop Hitting Yourself: Leveraging helpful assistance as an attack vector}},
  author = {Dustin Venini},
  year = {2025},
  month = mar,
  note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/stop-hitting-yourself-leveraging-helpful-assistance-as-an-attack-vector-w5rt}},
  url = {https://apartresearch.com/sprints/projects/stop-hitting-yourself-leveraging-helpful-assistance-as-an-attack-vector-w5rt}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026