cs.CVMay 22, 2026

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images

Authors: Zhaonan LiKyle R. ChickeringBangzheng LiJacob DineenXiao YeZhikun XuShijie LuYuxi Huang+8 more

Organizations: Arizona State University · Luma AI · UC Davis

Abstract

A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties under transformation and transfer them to new scenes. We introduce VisAnalog, a controlled suite for this setting on natural images. Each example instantiates A ⁣: ⁣B::C ⁣:?A\!:\!B::C\!:\,?: images BB and a hidden target image DD are produced by applying the same deterministic transformation sequence to source images AA and CC. Given AA, BB, and CC, a model must answer a multiple-choice question about DD. The benchmark contains 617 human-validated questions spanning one- to four-step transformations such as zoom, quadrant swap, rotation, flip, and hue rotation. Across strong proprietary and open-source VLMs, end-to-end accuracy is substantially lower than oracle accuracy when DD is directly shown, and degrades sharply as transformation depth increases, while human performance remains near the ceiling. A program-conditioned evaluation further separates failures of relation inference from failures of transformation application, showing that inferring the visual relation from ABA \rightarrow B is the dominant bottleneck, with additional application errors emerging on harder multi-step cases. The dataset is publicly available at https://huggingface.co/datasets/zli99/VisAnalog.

Explore similar work

CardsList
  1. Show Me Examples: Inferring Visual Concepts from Image Sets

    Jul 2, 2026Nick Stracke, Kolja Bauer, Stefan Andreas Baumann +3Visual ContextVisual Reasoning