Out of Sight, Out of Mind: Image Preferences in Vision-Language Models
Swante Scholz
Do vision-language models have preferences about what they look at? We measure stated and revealed preference over ten images — five categories × two exemplars, including noise and solid colour as controls — in four models from four labs.
They do. Stated ratings predict revealed choice in every model (ρ = 0.57–0.98), and the degenerate categories take 0–4% of 200 choices. Not a complexity effect: noise is the most complex stimulus, and among the least chosen.
Given repeated choices, models tour: category shares stay near-uniform until every image has been seen. That drive depends on the model retaining its own prior turns. Remove them — a routine context-management operation — and three of four collapse from touring nearly all ten images to revisiting one or two. What a preference measurement finds therefore depends on how the conversation is structured, not on the model alone.
This was a well thought out project. I really liked the direction they went in for image preferences and context compression. This was also unusually well written and communicated! finding that context compression changes preferences completely. Measuring the difference between revealed choice and stated preference is also rigorous. I think the research problem was done well but the implications for the field are missing. Why does it matter whether models appear to prefer one image over another? If we don't believe this finding impacts model welfare which is a big leap and not proven here the work does not have much implications. I think this person has clear skills in AI research and should continue but research taste could use some development.
This submission asks whether vision-language models have preferences about what they look at, measuring stated ratings and repeated choices over the same small image set in four models from four labs, and reporting both that self-report tracks choice and that removing a model's own prior turns from its context collapses exploration into repetition. The process around the result is the strongest part of the work: predictions were registered before any data was collected, several discarded designs are reported together with the measurements that killed them, and the call-level record is public. The redaction condition would repay one more cell, because it currently changes three things at once, removing the assistant turns while also adding a paragraph to the system prompt describing the redaction study and a parenthetical to every user turn restating that the reasoning is missing, so a condition that removes only the turns and leaves the prompt wording identical to the full-history arm would separate memory loss from a response to being told about the manipulation. The correlation between stated rating and revealed choice also needs its per-model values, its unit of analysis and an interval reported alongside it, since at ten items the low end of the reported range does not exclude zero and can be produced by the two degenerate stimuli sitting last on both measures while the photographs are ranked independently.
Cite this work
@misc {
title={
(HckPrj) Out of Sight, Out of Mind: Image Preferences in Vision-Language Models
},
author={
Swante Scholz
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


