cs.CVSep 7, 2026

Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation

Authors: Shangzhe DiZhaokai WangWeidi Xie

Abstract

Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting. We introduce ProbeGen, a benchmark for zero-shot generative perception that casts monocular depth estimation, referring/reasoning segmentation, and object counting as conditional generation tasks specified through text prompts, and compares 20 models in total---including proprietary and open-weight image generators, specialist perception models, and MLLMs---across 11 published benchmarks. We observe that pretrained image generators show measurable zero-shot perceptual competence, but with a clear trade-off: specialist models remain stronger for in-distribution accuracy and efficiency, while generative models are often more robust under distribution shift and better at compositional semantic reasoning. We hope this study helps establish zero-shot generative perception as a meaningful research direction and provides a useful foundation for future work at the intersection of visual generation and understanding.

Explore similar work

CardsList
  1. Image Generators are Generalist Vision Learners

    Apr 22, 2026Valentin Gabeur, Shangbang Long, Songyou Peng +22Visual GenerationImage Generation