Utilitarian Decision-Making in Models - Evaluation and Steering
Adam Newgas, Sinem Erisken , Pandelis Mouratoglou · Team Byte To Brain
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering can alter this preference, and the ability of the Llama 70B model to impersonate other people’s opinions. Results are validated against a dataset of 10,000 human responses. We find that (1) Llama does not accurately capture human moral values, (2) OUS offers better interpretations than current feature labels, and (3) Llama fails to predict the demographics of human values.

Reviews
This is very well executed and presented research. The comparison of model values vs the human response KDE is interesting, but my favourite plot is figure 2 - it's very surprising how different features have remarkably different trajectories through the moral landscape. It's surprising that most features actually appear to avoid the modal human, and only a single feature actually steers the model in that direction. It's unfortunate that the OUS has so few questions and is so sensitive (e.g. the difference between models being entirely accounted for by question IH2).
This is pretty cool and well thought out!
I found this project really interesting! It is surprising how poorly the LLMs seemed to model human moral intuitions, even when steered. This is well-written and well-presented!
This is a great project! Very well done. I think it's good it has a lower IH than humans and that this IH is less variable as well. It's very interesting to see these underlying patterns and the comparison to human baselines is very interesting. It seems like there are clear pathways forward to studying the moral impact of irrelevant features, such as the elephant features you mention. You may also run the model at a higher temperature and get the density function for each model to compare with human values. There is of course an implicit assumption that utilitarianism is opposed to deontology in the methodological design and it would be great to check if these patterns shift under e.g. legalistic scenarios vs. moral dilemmas (https://paperswithcode.com/dataset/ethics-1).
Cite this project
@misc{newgas2024utilitarian,
title = {{Utilitarian Decision-Making in Models - Evaluation and Steering}},
author = {Adam Newgas and Sinem Erisken and Pandelis Mouratoglou},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/utilitarian-decision-making-in-models-evaluation-and-steering}},
url = {https://apartresearch.com/sprints/projects/utilitarian-decision-making-in-models-evaluation-and-steering}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on …
- View project: Steering Swiftly to Safety with Sparse Autoencoders
Steering Swiftly to Safety with Sparse Autoencoders
Explaining_Polysemantic_Feature_Learning
We explore using SAEs for unlearning dangerous capabilities in a cheaper and more interpretable way.