Skip to content
Sprint projectNov 3, 2025Lausanne

Empirical Measurements of Technique Effectiveness Across Model Sizes

Arthur Wuhrmann, Ines Altemir Marinas, Kyuhee Kim · Team Safe AI Lausanne

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Empirical Measurements of Technique Effectiveness Across Model Sizes

Code (opens in new tab)
Share

We estimated the evolution of AI Safety techniques and demonstrated evidence of predictive power. We emphasize on the necessity of evaluating safety techniques across different model sizes to ensure their robustness and predictive power.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. * I like the use of the safetywashing paper to decrease the amount that the assessments you use are actually measuring capabilities.

    * While TQA is a somewhat easy dataset to work with, it has severe limitations and really doesn't measure "truthfullness" at all (just a side note from someone who has worked extensively with the dataset). The fact that you found differing results from Ren et al. on this also backs up that we shouldn't be using the dataset anymore; it's incredibly confusing to work with due to it's non-standard design, meaning that almost no one actually uses the same assessment methodology. I'll stop bashing on TQA now...

    * I really appreciate the breadth of data shared in the submission, but this makes it quite difficult to parse 'takeaways' from the different figures (not to mention the figure sizes are far too small). It would be beneficial to disentangle different aspects of the data in a more thorough study, but I understand that this would require replicated the studies used, which isn't feasible for various reasons.

    * Another aspect which would be interesting to see in parallel with what is presented is what the cost of the safety intervention looks like, and how that scales with model size.

    * I appreciate that you point out a key failure mode of AI safety benchmarking works! This is quite important, and the more people we have who know why comprehensive assessment and reporting is valuable, the better this will be in the future.

    Read full reviewShow less
  2. The project focuses on understanding how safety intervention techniques scale with model size. First, they look at the performance of the models on the safety benchmarks. Then, they apply safety intervention techniques and see how much they improve on the benchmarks, finding that the performance gains of different techniques differ depending on the benchmark.

    I generally think that this is pretty interesting work. My main concern here would be how it is probably less likely for models to continuously get bigger in model size, so an x-axis of capabilities score (similar to the safety washing paper) may be more appropriate for better anchoring the abilities of the paper.

    However, I am also concerned about the lack of details in terms of the implementation of the safety interventions: most of them are likely to be sensitive to theirhyperparameters, and I would not be sure how thorough the interventions were implemented. I think this is fine given as a weekend hackathon project, and would encourage the authors to further look into how the important determining the right hyperparameters of these techniques are for this project.

    Lastly, it may also be valuable for the authors to focus on one or two specific interventions they think have promise in scaling.

    Read full reviewShow less

Cite this project

@misc{wuhrmann2025empirical,
  title = {{Empirical Measurements of Technique Effectiveness Across Model Sizes}},
  author = {Arthur Wuhrmann and Ines Altemir Marinas and Kyuhee Kim},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/empirical-measurements-of-technique-effectiveness-across-model-sizes-swyd}},
  url = {https://apartresearch.com/sprints/projects/empirical-measurements-of-technique-effectiveness-across-model-sizes-swyd}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026