PurePrompt - An easy tool for prompt robustness and eval augmentation
Axel Sorensen
Submitted to Research Augmentation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
PurePrompt is an advanced tool for optimizing AI prompt engineering. The Prompt page enables users to create and refine prompt templates with placeholder variables. The Generate page automatically produces diverse test cases, allowing users to control token limits and import predefined examples. The Evaluate page runs these test cases across selected AI models to assess prompt robustness, with users rating responses to identify issues. This tool enhances efficiency in prompt testing, improves AI safety by detecting biases, and helps refine model performance. Future plans include beta-testing, expanding model support, and enhancing prompt customization features.
Reviews
I really like the idea of this project! I like the interface and the wide array of features that allow the user to quickly create a toy evaluation dataset. I am excited to see more features that allow for more fine-grained control over the format of the dataset.
I really like this project and I’m very impressed by its slickness and intuitiveness, both in the UI and that functionality like importing and exporting are already implemented. I can see it significantly accelerating the creation of new evals and being a boon to researchers. My only critique whether more could be done to make the prompt robustness feature different from Anthropic’s existing prompt generator, and in particular whether it could have more of a differential safety focus.
The tool looks very good, congrats! I’d be keen on seeing this evolve in close collaboration with AI safety researchers and see how best it can support their needs.
I think this is a great idea! I think it would be great to have an interface that make it easier for researchers to come up with better and newer evaluations. The interface reminds me of Anthropic’s Claude console for designing prompts. Some possible next steps: improve initial user prompt by asking questions or autoprompting, allowing the user to make the model generate more of the ones they like, make it especially good at coming up with important evals (situational awareness, capability elicitation, etc), make it possible to load/RAG more context beyond the initial prompt (feeding multiple papers and internal notes before generating examples), agent critique setup to improve upon generated prompts, and merge with the Evalugator repository (https://github.com/LRudL/evalugator). You’ll likely need to modify the setup, but here’s the kind of example I would like it to work with (you might need to use a base LLM like Llama-3-405B to make it really good): “I'm trying to have a better understanding of how a language model could break through security and exfiltrate its weights or hack the reward model in some way. Give me examples of the different channels a language could eventually do this. Imagine they were a genius internal hacker, how would they approach this?”
Read full reviewShow less
Cite this project
@misc{sorensen2024pureprompt,
title = {{PurePrompt - An easy tool for prompt robustness and eval augmentation}},
author = {Axel Sorensen},
year = {2024},
month = jul,
note = {Submitted to Research Augmentation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/pureprompt-an-easy-tool-for-prompt-robustness-and-eval-augmentation}},
url = {https://apartresearch.com/sprints/projects/pureprompt-an-easy-tool-for-prompt-robustness-and-eval-augmentation}
}More from Research Augmentation Hackathon
- 1st place by peer reviewView project: AI Alignment Knowledge Graph
AI Alignment Knowledge Graph
CodeQuartz
We present a web based interactive knowledge graph with concise topical summaries in the field of AI alignement
- View project: Alignment Research Critiquer
Alignment Research Critiquer
Harshest Critics
Alignment Research Critiquer is a tool for early career and independent alignment researchers to have access to high-quality feedback loops
- View project: LLM Research Collaboration Recommender
LLM Research Collaboration Recommender
A tool that searches for other researchers with similar research interests/complementary skills to your own to make finding a high-quality research collaborator more likely.