cs.CVFeb 16, 2026

Scalable next-scale autoregression for medical image generation across anatomical regions

Authors: Zhicheng He, Yunpeng Zhao, Junde Wu, Ziwei Niu, Ziyue Wang, Bohan Li, Zijun Li, Lanfen Lin, +2 more

Organizations: National University of Singapore · University of Oxford · Zhejiang University · Shanghai Jiao Tong University

Abstract

Autoregressive pretraining has been key to the scalability of large language models, yet medical generative foundation models remain predominantly based on diffusion. Here we introduce MedVAR, the first foundation model for all-round medical image generation through autoregressive training, and find it offers improved generation quality and efficiency, stronger scalability, and broader adaptability to downstream clinical tasks than diffusion-based models. A tokenizer trained on medical images and separate semantic and structural controls enable generation across six anatomical regions in computed tomography and magnetic resonance imaging. Trained on 438,905 slices from 40 datasets, including seven internal clinical centres, MedVAR generates images 12-19 times faster than 100-step diffusion baselines. Generation quality improves with model size. Pretraining on generated images improves seven-centre hepatocellular carcinoma segmentation Dice from 0.615 to 0.632, while reconstruction using MedVAR images approaches the volumetric segmentation performance of fully sampled volumes. Membership inference reaches 3.54% sensitivity at a 1% false-positive rate, while copy detection performs near chance. These findings establish next-scale autoregression as a scalable and versatile approach to medical image generation and downstream analysis.

Figures & tables

Explore similar work

CardsList
  1. Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling

    Jul 10, 2026Chicago Y. Park, Jialin Mao, Xiaojian Xu +3Autoregressive Image GenerationDense Prediction

  2. Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling

    Sep 12, 2026Meimingwei Li, Stefan Andreas Baumann, Felix Krause +1Visual Autoregressive ModelsLogit Lens

  3. TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model

    Jul 15, 2026Zhenkai Zhang, Krista A. Ehinger, Tom DrummondMedical Image Generation3D Generative Models