Skip to content
Sprint projectJul 29, 2024

PurePrompt - An easy tool for prompt robustness and eval augmentation

Axel Sorensen

Submitted to Research Augmentation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: PurePrompt - An easy tool for prompt robustness and eval augmentation

Code (opens in new tab)More on pureprompt.vercel.app (opens in new tab)
Share

PurePrompt is an advanced tool for optimizing AI prompt engineering. The Prompt page enables users to create and refine prompt templates with placeholder variables. The Generate page automatically produces diverse test cases, allowing users to control token limits and import predefined examples. The Evaluate page runs these test cases across selected AI models to assess prompt robustness, with users rating responses to identify issues. This tool enhances efficiency in prompt testing, improves AI safety by detecting biases, and helps refine model performance. Future plans include beta-testing, expanding model support, and enhancing prompt customization features.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. I really like the idea of this project! I like the interface and the wide array of features that allow the user to quickly create a toy evaluation dataset. I am excited to see more features that allow for more fine-grained control over the format of the dataset.

  2. I really like this project and I’m very impressed by its slickness and intuitiveness, both in the UI and that functionality like importing and exporting are already implemented. I can see it significantly accelerating the creation of new evals and being a boon to researchers. My only critique whether more could be done to make the prompt robustness feature different from Anthropic’s existing prompt generator, and in particular whether it could have more of a differential safety focus.

  3. The tool looks very good, congrats! I’d be keen on seeing this evolve in close collaboration with AI safety researchers and see how best it can support their needs.

  4. I think this is a great idea! I think it would be great to have an interface that make it easier for researchers to come up with better and newer evaluations. The interface reminds me of Anthropic’s Claude console for designing prompts. Some possible next steps: improve initial user prompt by asking questions or autoprompting, allowing the user to make the model generate more of the ones they like, make it especially good at coming up with important evals (situational awareness, capability elicitation, etc), make it possible to load/RAG more context beyond the initial prompt (feeding multiple papers and internal notes before generating examples), agent critique setup to improve upon generated prompts, and merge with the Evalugator repository (https://github.com/LRudL/evalugator). You’ll likely need to modify the setup, but here’s the kind of example I would like it to work with (you might need to use a base LLM like Llama-3-405B to make it really good): “I'm trying to have a better understanding of how a language model could break through security and exfiltrate its weights or hack the reward model in some way. Give me examples of the different channels a language could eventually do this. Imagine they were a genius internal hacker, how would they approach this?”

    Read full reviewShow less

Cite this project

@misc{sorensen2024pureprompt,
  title = {{PurePrompt - An easy tool for prompt robustness and eval augmentation}},
  author = {Axel Sorensen},
  year = {2024},
  month = jul,
  note = {Submitted to Research Augmentation Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/pureprompt-an-easy-tool-for-prompt-robustness-and-eval-augmentation}},
  url = {https://apartresearch.com/sprints/projects/pureprompt-an-easy-tool-for-prompt-robustness-and-eval-augmentation}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026