Data Massager
Giles Edkins, Ayush Jain · Team RolyPoly
Submitted to Research Augmentation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A VSCode plugin for helping the creation of Q&A datasets used for the evaluation of LLMs capabilities and alignment.
Reviews
Great project! I really like the wide array of features that are well thought out from the perspective of the needs of the user and the integration with VS Code! I would be excited to see features that help automate the data generation and curation pipeline if the steps are similar to what the user has done in the past.
I like this project a lot, and I can see it really shining when using large models to create datasets for novel benchmarks for smaller models. I like how intuitive the UI in VSCode is, and that the prospective dataset creator doesn’t have to leave their IDE to use the tool. As a further step, perhaps you could consider prompting the LLM to give a confidence score for its answer - humans could then review all answers below a certain confidence score, or simply choose to remove them all from the dataset.
This project seems very promising and the writeup is very helpful for understanding the current scope, strengths and weaknesses. For me the most important question to address here is how to measure the reliability of the different operations on the fly and how to improve it. I think it will be very hard (and for some datasets maybe impossible in the near-term) to build a faultless system, so the core question is how to get a tool that is aware of its own limitations and lets the human researchers work around those limitations.
I think the general idea for this project is good. There probably should have been more focus on figuring out how to make it uniquely good at coming up with new alignment evaluations (is the model being deceptive? is it sandbagging? does it have situational awareness?) and trying to make it especially good at helping the researcher coming it with new examples. Overall, I like the different features and think there is some shape of this kind of project that could make it significantly easier to generate higher quality evals.
Cite this project
@misc{edkins2024data,
title = {{Data Massager}},
author = {Giles Edkins and Ayush Jain},
year = {2024},
month = jul,
note = {Submitted to Research Augmentation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/data-massager}},
url = {https://apartresearch.com/sprints/projects/data-massager}
}More from Research Augmentation Hackathon
- 1st place by peer reviewView project: AI Alignment Knowledge Graph
AI Alignment Knowledge Graph
CodeQuartz
We present a web based interactive knowledge graph with concise topical summaries in the field of AI alignement
- View project: Alignment Research Critiquer
Alignment Research Critiquer
Harshest Critics
Alignment Research Critiquer is a tool for early career and independent alignment researchers to have access to high-quality feedback loops
- View project: PurePrompt - An easy tool for prompt robustness and eval augmentation
PurePrompt - An easy tool for prompt robustness and eval augmentation
PurePrompt is an advanced tool for optimizing AI prompt engineering. The Prompt page enables users to create and refine prompt templates with placeholder variables. The Generate page automatically produces diverse test …