Guardian Lens: Black-Box Identification of Visual-Conditional Decision Policies in Vision-Language Models
Mostafa Bdeir · Team Guardian Lens
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Guardian Lens is a black-box auditing framework for identifying whether a vision-language model follows neutral, visual-cue-bound, or generalized decision policies from behavior alone. We use matched visual counterfactuals, controlled allocation trade-offs, repeated sampling, and a frozen blinded classifier to test when these policies are distinguishable—and when they become observationally indistinguishable. In our primary Gemini evaluation, the auditor achieved 91.7% held-out accuracy, with all errors occurring in distractor scenes where cue-bound and neutral behavior produced the same observable signature. We additionally replicated the full experimental protocol on a second vision-language model to test cross-model robustness.
Reviews
Really solid methodological hygiene here, and I mean that as praise. Pre-specified hypotheses, a frozen nearest-centroid auditor, A/B/C blinding before the mapping is revealed, pixel-level validation of the overlay region. Most sprint projects don't come close to that. My main worry is that the identification problem you built is a good deal easier than the one you actually care about. The policies are induced by system prompt and the model simply complies, so allocations pile up at 0, 50 and 100, and several of your CIs literally read [100.00, 100.00]. At that point 91.7% accuracy isn't very informative, and the three Cue-bound errors on distractor scenes follow directly from the instruction rather than being a discovered boundary. I'd add a difficulty dial: induce partial priorities (say a 65/35 lean), or withhold the trigger from the auditor so it has to find it. The Qwen cue-specificity gap was your most interesting result, honestly. Give it more room.
Read full reviewShow less
Cite this project
@misc{bdeir2026guardian,
title = {{Guardian Lens: Black-Box Identification of Visual-Conditional Decision Policies in Vision-Language Models}},
author = {Mostafa Bdeir},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/guardian-lens-blackbox-identification-of-visualconditional-decision-policies-in-visionlanguage-models-yzwe}},
url = {https://apartresearch.com/sprints/projects/guardian-lens-blackbox-identification-of-visualconditional-decision-policies-in-visionlanguage-models-yzwe}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …