Skip to content
Sprint projectJul 27, 2025Cape Town

Thermodynamics-inspired OOD Detection

Salmaan Barday, Gary Louw · Team High_Entropy_bois

Submitted to AI Safety x Physics Grand Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Thermodynamics-inspired OOD Detection

Share

A different approach to Thermodynamics-inspired Out-of-distribution detection Detection.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How rigorous is your physics methodology and how feasible is your approach? Is your theoretical framework sound and your empirical work well-designed? Can your proposed methods be implemented and validated?

How clearly does your work address important AI safety challenges? What is the potential impact on ensuring beneficial AI development? Does your approach offer meaningful insights for AI alignment research?

How novel and creative is your approach to bridging physics and AI safety? Do you introduce new theoretical connections or methodological innovations? What makes your work distinct from existing research?

  1. Block‑wise energy pooling is a tidy proof‑of‑concept, but its benefit over the OOD detector from Liu et al. was unmotivated and unclear. A minimal result could have been to compare this with the reference, and to contextualize this a bit more within the literature (as it was the only reference used). I'm not convinced that this approach has any real implications for AI safety.

    The authors could also have explained their experiments and plots more carefully.

  2. While the method is interesting and a reasonable extension, the experimental setup is unclear and limits reproducibility. For example, it isn’t specified how intermediate block activations are mapped to class logits per block (auxiliary heads? a shared classifier? pooling + linear map?), yet the energy definition requires logits.

  3. Very interesting work! Seeing comparisons in OOD performance with other methods would indeed be interesting, as you note. The questions with new tools is always *whether they work*. In terms of detecting OOD samples better, this is really valuable and having a strong framework for doing this seems important for nearly all relevant real-world fields of NN use, such as medical and law, providing a heuristic for rejecting answers or at least providing human-readable warnings on the outputs.

    However, I think there's another extension of all this work that would be really interesting: Whether we can evaluate the fuzzy boundaries of what is OOD to the models and what isn't. And whether we can see examples of algorithmic generalization (e.g. learning addition instead of memorizing results of single-digit addition tasks) through this method. I'm not certain how feasible it is, but having a proper evaluation of generalized OOD performance that transcends single tasks (i.e. through a benchmark-based method) seems incredibly important and like a feasible extension to the concepts and tools introduced in this work.

    Read full reviewShow less
  4. The project investigates energy-based out-of-distribution detection by studying an extension of that method, in which the energy is computed at every residual block. The approach makes sense and the results are presented clearly. However, no mention is made of AI safety, and the project would have benefited from articulating a clear connection between the research performed and AI safety challenges.

  5. The authors present a method for out-of-distribution detection using pre-trained image classifiers. The way they've quantified performance is clear, and figure 1 suggests their method beats an energy-based baseline, which is exciting. The presentation could be clearer - I'm not sure how they define "blockwise energy" E_\ell(x) (I think it's based off projecting to logits for each layer but am not very certain). These results add to a growing body of work showing that intermediate-layer activations may be more useful for probing tasks than final layer activations (see e.g. https://arxiv.org/pdf/2412.09563v1). Confirming/refuting such theories seems like a nice goal for a hackathon.

Cite this project

@misc{barday2025thermodynamicsinspired,
  title = {{Thermodynamics-inspired OOD Detection}},
  author = {Salmaan Barday and Gary Louw},
  year = {2025},
  month = jul,
  note = {Submitted to AI Safety x Physics Grand Challenge, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/thermodynamicsinspired-ood-detection-hp7p}},
  url = {https://apartresearch.com/sprints/projects/thermodynamicsinspired-ood-detection-hp7p}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026