Skip to content
Sprint projectNov 25, 2024

Utilitarian Decision-Making in Models - Evaluation and Steering

Adam Newgas, Sinem Erisken , Pandelis Mouratoglou · Team Byte To Brain

Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Utilitarian Decision-Making in Models - Evaluation and Steering

Code (opens in new tab)
Share

We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering can alter this preference, and the ability of the Llama 70B model to impersonate other people’s opinions. Results are validated against a dataset of 10,000 human responses. We find that (1) Llama does not accurately capture human moral values, (2) OUS offers better interpretations than current feature labels, and (3) Llama fails to predict the demographics of human values.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project contribute to the field of mechanistic interpretability? Does it provide new insights into understanding or steering AI model behavior? How well does it move us towards reprogramming AI models? How original and innovative is the approach?

How important is the contribution to advancing the field of AI safety? Do we expect the results to generalize beyond the specific case(s) presented in the submission? Does the approach introduce new safety mechanisms or enhance existing ones in innovative ways?

How well is the project executed from a technical standpoint?, Is the code well-structured, documented, and reproducible?, How effectively does it utilize Goodfire's SDK/API and other provided resources?, How clearly and effectively is the research presented in the paper?, Quality of visualizations and demos (if applicable), Clarity of methodology explanation and results interpretation

  1. This is very well executed and presented research. The comparison of model values vs the human response KDE is interesting, but my favourite plot is figure 2 - it's very surprising how different features have remarkably different trajectories through the moral landscape. It's surprising that most features actually appear to avoid the modal human, and only a single feature actually steers the model in that direction. It's unfortunate that the OUS has so few questions and is so sensitive (e.g. the difference between models being entirely accounted for by question IH2).

  2. This is pretty cool and well thought out!

  3. I found this project really interesting! It is surprising how poorly the LLMs seemed to model human moral intuitions, even when steered. This is well-written and well-presented!

  4. This is a great project! Very well done. I think it's good it has a lower IH than humans and that this IH is less variable as well. It's very interesting to see these underlying patterns and the comparison to human baselines is very interesting. It seems like there are clear pathways forward to studying the moral impact of irrelevant features, such as the elephant features you mention. You may also run the model at a higher temperature and get the density function for each model to compare with human values. There is of course an implicit assumption that utilitarianism is opposed to deontology in the methodological design and it would be great to check if these patterns shift under e.g. legalistic scenarios vs. moral dilemmas (https://paperswithcode.com/dataset/ethics-1).

Cite this project

@misc{newgas2024utilitarian,
  title = {{Utilitarian Decision-Making in Models - Evaluation and Steering}},
  author = {Adam Newgas and Sinem Erisken and Pandelis Mouratoglou},
  year = {2024},
  month = nov,
  note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/utilitarian-decision-making-in-models-evaluation-and-steering}},
  url = {https://apartresearch.com/sprints/projects/utilitarian-decision-making-in-models-evaluation-and-steering}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026