Skip to content
Sprint projectMar 30, 2025Singapore

Safety Metric and Prompt Engineering for Red Team

Zhu Liang · Team Paradite

Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Safety Metric and Prompt Engineering for Red Team

Share

In the ControlArena environment, we have a default attack policy that is very basic. We want to explore how to improve the attack policy using basic prompt engineering techniques.

We also realized that the security metric evaluation in the environment is not closely following the foundational paper (Greenblatt et al., 2023), so we made a change to more closely approximate the security metric in the paper.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{liang2025safety,
  title = {{Safety Metric and Prompt Engineering for Red Team}},
  author = {Zhu Liang},
  year = {2025},
  month = mar,
  note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/safety-metric-and-prompt-engineering-for-red-team-q2d6}},
  url = {https://apartresearch.com/sprints/projects/safety-metric-and-prompt-engineering-for-red-team-q2d6}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026