Autoregressive pretraining has been key to the scalability of large language models, yet medical generative foundation models remain predominantly based on diffusion. Here we introduce MedVAR, the first foundation model for all-round medical image generation through autoregressive training, and find it offers improved generation quality and efficiency, stronger scalability, and broader adaptability to downstream clinical tasks than diffusion-based models. A tokenizer trained on medical images and separate semantic and structural controls enable generation across six anatomical regions in computed tomography and magnetic resonance imaging. Trained on 438,905 slices from 40 datasets, including seven internal clinical centres, MedVAR generates images 12-19 times faster than 100-step diffusion baselines. Generation quality improves with model size. Pretraining on generated images improves seven-centre hepatocellular carcinoma segmentation Dice from 0.615 to 0.632, while reconstruction using MedVAR images approaches the volumetric segmentation performance of fully sampled volumes. Membership inference reaches 3.54% sensitivity at a 1% false-positive rate, while copy detection performs near chance. These findings establish next-scale autoregression as a scalable and versatile approach to medical image generation and downstream analysis.
Figures & tables
Figure 1: A next-scale autoregressive generator across anatomical regions. a , Anatomy and modality coverage. b , Distribution of 482,815 images from 40 datasets. c , Conditioning, generation and evaluation workflow. d , Seven-centre HCC MRI cohort; labels give volume counts and red contours mark tumours. e , FID comparisons within anatomical groups and a descriptive measure of quality and efficiency normalized to depth-30 MedVAR (Methods).
Figure 2: Generation quality, efficiency and diversity across MedVAR model sizes and generator families. a , Matched examples. b , Model depths. c , FID ranks within anatomy–modality groups using MedVAR with semantic conditions alone. d , FID versus latency; MedVAR depths 16 and 20 approximately match diffusion models at 330 and 600 million parameters, respectively. Diffusion curves span 10–100 steps; paired MedVAR points use semantic conditions alone or both semantic and structural conditions. e , Fidelity across generator families. f , Condition recognizability and CAS; the dashed line marks random balanced accuracy ( 1/9 ). g , Diversity under fixed conditions. N.A., not available; error bars, 95% bootstrap CIs.
Figure 3: Medical tokenization and conditioning support explicit control. a , Natural and medical tokenizer comparison. b , Progressive conditioning. c , Conditioning ablation. d , Ablations of structural routing, the scale gate and the auxiliary structure head; MedVAR routes token maps of sizes 52 , 62 , 82 and 102 . e , Effect of the weight applied to structural conditions. f , CFG and truncation ablations. Filled and hatched bars denote generation with semantic conditions alone and with both semantic and structural conditions, respectively; arrows indicate lower is better.
Figure 4: Generated images support segmentation and reconstruction of missing slices. a , HCC segmentation and reconstruction of missing slices. b , Dice gain for each centre over pretraining on real images; points denote matched seeds. c , Real images, images generated with centre labels and unconditional MedVAR images. d , PCA and t-SNE of RadImageNet features. e , Accuracy of centre prediction and maximum mean discrepancy (MMD); the dashed line is random. f , Matched CT reconstructions (top), pairwise signed differences for Linear–U-Net, U-Net–MedVAR and MedVAR–Real on a common scale (middle), and mean CT and MRI PSNR across volumes (bottom). Linear interpolation is deterministic, the U-Net is an independent residual baseline, and circles denote the three training seeds for the system using MedVAR. g , Paired 3D Dice distributions for linear interpolation, the residual U-Net, reconstruction using MedVAR and fully sampled input; boxes show medians and interquartile ranges, whiskers extend to 1.5 times the interquartile range, and points denote volumes. Arrows indicate better performance.
Figure 5: Privacy analyses and effects of data source labels. a , White-box membership inference by ROC AUC and TPR at 1% FPR; error bars are 95% bootstrap CIs. b , ROC curves at low FPR; dashed lines mark 1% FPR and random performance. c , Nearest-neighbour copy detection for three sizes of the generated set. Diamonds and bars show mean ± s.d. across encoders; circles show individual encoders. d , Generated images and nearest training and test images with NCC, SSIM and difference maps. e , Stress test with and without labels identifying the data source under simulated contamination; positive interaction values indicate attenuation when the label is provided.
Figure 6: MedVAR architecture and semantic and structural conditioning. a , Coarse-to-fine token prediction with semantic and optional structural conditions; the auxiliary structure head is used only during training. b , Hierarchical semantic conditioning combines region and organ labels, then fuses this representation with imaging modality. A separately projected data source embedding is added at reduced weight before adaptive layer normalization. c , Structural attention routing embeds the organ mask and its boundary at each token-map scale. These features update the query and key projections of self-attention; the value projection is unchanged.
We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer. Our key insight is that traversing a single-scale latent grid with progressively denser strides naturally captures the transition from global structure to fine detail. This addresses two limitations of existing autoregressive models at once: the slow inference of raster-order autoregression, which DenseAR avoids by predicting multiple tokens in parallel, and the heavy cost of multi-scale approaches, which need long, multi-resolution token sequences to achieve coarse-to-fine prediction. Building on our efficient framework and the flexibility of autoregressive modeling, we further extend DenseAR to a unified model that handles multiple modalities and imaging tasks within a single backbone. We validate DenseAR on both medical and natural images. On multi-contrast brain MRI, a single DenseAR model unifies cross-modal translation, modality-conditioned generation, and tumor segmentation, while remaining competitive with task-specific methods. On ImageNet, DenseAR improves class-conditional generation quality (FID and IS) over both a single-grid baseline without stride ordering and a multi-scale tokenizer-based baseline.
Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless of backbone capacity -- a limitation of the decoding rule. Addressing this limitation, we introduce the Logit Refiner, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features. Adding only ~10% parameters and less than 5% of the base model's training compute, it plugs into any pretrained VAR checkpoint without retraining. Controlled ablations isolate joint intra-scale sampling -- rather than additional capacity or training -- as the critical ingredient. Across backbones from 310M to 2B parameters on class-conditional ImageNet 256x256, the refiner consistently improves generation quality, enabling a 1.1B-parameter model to surpass one twice its size. The approach further generalizes to text-to-image generation, confirming that the mean-field bottleneck persists across VAR variants and is effectively alleviated by our method. Project page: https://compvis.github.io/logit-refiner/
Meimingwei Li, Stefan Andreas Baumann, Felix Krause +1
We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting. Subsequently, it uses a triplane-aware cross-attention diffusion model to learn and integrate these features effectively. Furthermore, the features generated by the diffusion model can be rapidly transformed into 3D volumes using a pre-trained decoder module. Our experiments on three different scales of medical datasets, BrainTumour 128 x 128 x 128, Pancreas 256 x 256 x 256, and Colon 512 x 512 x 512, demonstrate outstanding results. We utilized MSE and SSIM to assess reconstruction quality and leveraged the Wasserstein Generative Adversarial Network (W-GAN) critic to assess generative quality. Comparisons with existing approaches show that our method gives better reconstruction and generation results than other encoder-decoder methods with similar-sized latent spaces.
Zhenkai Zhang, Krista A. Ehinger, Tom Drummond
School of Computing and Information Systems, The University of Melbourne · The University of Melbourne