Empirical Measurements of Technique Effectiveness Across Model Sizes
Arthur Wuhrmann, Ines Altemir Marinas, Kyuhee Kim · Team Safe AI Lausanne
Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We estimated the evolution of AI Safety techniques and demonstrated evidence of predictive power. We emphasize on the necessity of evaluating safety techniques across different model sizes to ensure their robustness and predictive power.
Reviews
* I like the use of the safetywashing paper to decrease the amount that the assessments you use are actually measuring capabilities.
* While TQA is a somewhat easy dataset to work with, it has severe limitations and really doesn't measure "truthfullness" at all (just a side note from someone who has worked extensively with the dataset). The fact that you found differing results from Ren et al. on this also backs up that we shouldn't be using the dataset anymore; it's incredibly confusing to work with due to it's non-standard design, meaning that almost no one actually uses the same assessment methodology. I'll stop bashing on TQA now...
* I really appreciate the breadth of data shared in the submission, but this makes it quite difficult to parse 'takeaways' from the different figures (not to mention the figure sizes are far too small). It would be beneficial to disentangle different aspects of the data in a more thorough study, but I understand that this would require replicated the studies used, which isn't feasible for various reasons.
* Another aspect which would be interesting to see in parallel with what is presented is what the cost of the safety intervention looks like, and how that scales with model size.
* I appreciate that you point out a key failure mode of AI safety benchmarking works! This is quite important, and the more people we have who know why comprehensive assessment and reporting is valuable, the better this will be in the future.
Read full reviewShow less
The project focuses on understanding how safety intervention techniques scale with model size. First, they look at the performance of the models on the safety benchmarks. Then, they apply safety intervention techniques and see how much they improve on the benchmarks, finding that the performance gains of different techniques differ depending on the benchmark.
I generally think that this is pretty interesting work. My main concern here would be how it is probably less likely for models to continuously get bigger in model size, so an x-axis of capabilities score (similar to the safety washing paper) may be more appropriate for better anchoring the abilities of the paper.
However, I am also concerned about the lack of details in terms of the implementation of the safety interventions: most of them are likely to be sensitive to theirhyperparameters, and I would not be sure how thorough the interventions were implemented. I think this is fine given as a weekend hackathon project, and would encourage the authors to further look into how the important determining the right hyperparameters of these techniques are for this project.
Lastly, it may also be valuable for the authors to focus on one or two specific interventions they think have promise in scaling.
Read full reviewShow less
Cite this project
@misc{wuhrmann2025empirical,
title = {{Empirical Measurements of Technique Effectiveness Across Model Sizes}},
author = {Arthur Wuhrmann and Ines Altemir Marinas and Kyuhee Kim},
year = {2025},
month = nov,
note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/empirical-measurements-of-technique-effectiveness-across-model-sizes-swyd}},
url = {https://apartresearch.com/sprints/projects/empirical-measurements-of-technique-effectiveness-across-model-sizes-swyd}
}More from The AI Forecasting Hackathon
- View project: System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics BCN
A Game theoretic / System Dynamics model of the race dynamics of the US, China, and EU, as a follow up to the Armstrong et al. (2016) paper “Racing to the Precipice”. We find preliminary results where knowledge of …
- View project: ExogenousAI
ExogenousAI
Fibonacci
Current AI capability forecasting methodologies, including EpochAI's Direct Approach and Biological Anchors framework, primarily rely on internal metrics such as training compute and scaling laws while assuming stable …
- View project: AI Incidents Forecasting
AI Incidents Forecasting
KLACE
This research develops a framework for forecasting AI incidents to help predict future risks. We have developed two models that forecasts incidents which include calibrated 90% prediction intervals with backtests. These …