Skip to content
Sprint projectAug 16, 2026Zurich

Out of Sight, Out of Mind: Image Preferences in Vision-Language Models

Swante Scholz · Team Swante

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Out of Sight, Out of Mind: Image Preferences in Vision-Language Models

Code (opens in new tab)
Share

Do vision-language models have preferences about what they look at? We measure stated and revealed preference over ten images — five categories × two exemplars, including noise and solid colour as controls — in four models from four labs.

They do. Stated ratings predict revealed choice in every model (ρ = 0.57–0.98), and the degenerate categories take 0–4% of 200 choices. Not a complexity effect: noise is the most complex stimulus, and among the least chosen.

Given repeated choices, models tour: category shares stay near-uniform until every image has been seen. That drive depends on the model retaining its own prior turns. Remove them — a routine context-management operation — and three of four collapse from touring nearly all ten images to revisiting one or two. What a preference measurement finds therefore depends on how the conversation is structured, not on the model alone.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This was a well thought out project. I really liked the direction they went in for image preferences and context compression. This was also unusually well written and communicated! finding that context compression changes preferences completely. Measuring the difference between revealed choice and stated preference is also rigorous. I think the research problem was done well but the implications for the field are missing. Why does it matter whether models appear to prefer one image over another? If we don't believe this finding impacts model welfare which is a big leap and not proven here the work does not have much implications. I think this person has clear skills in AI research and should continue but research taste could use some development.

  2. This submission asks whether vision-language models have preferences about what they look at, measuring stated ratings and repeated choices over the same small image set in four models from four labs, and reporting both that self-report tracks choice and that removing a model's own prior turns from its context collapses exploration into repetition. The process around the result is the strongest part of the work: predictions were registered before any data was collected, several discarded designs are reported together with the measurements that killed them, and the call-level record is public. The redaction condition would repay one more cell, because it currently changes three things at once, removing the assistant turns while also adding a paragraph to the system prompt describing the redaction study and a parenthetical to every user turn restating that the reasoning is missing, so a condition that removes only the turns and leaves the prompt wording identical to the full-history arm would separate memory loss from a response to being told about the manipulation. The correlation between stated rating and revealed choice also needs its per-model values, its unit of analysis and an interval reported alongside it, since at ten items the low end of the reported range does not exclude zero and can be produced by the two degenerate stimuli sitting last on both measures while the photographs are ranked independently.

    Read full reviewShow less

Cite this project

@misc{scholz2026out,
  title = {{Out of Sight, Out of Mind: Image Preferences in Vision-Language Models}},
  author = {Swante Scholz},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/out-of-sight-out-of-mind-image-preferences-in-visionlanguage-models-7v4c}},
  url = {https://apartresearch.com/sprints/projects/out-of-sight-out-of-mind-image-preferences-in-visionlanguage-models-7v4c}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026