Autoregressive pretraining has been key to the scalability of large language models, yet medical generative foundation models remain predominantly based on diffusion. Here we introduce MedVAR, the first foundation model for all-round medical image generation through autoregressive training, and find it offers improved generation quality and efficiency, stronger scalability, and broader adaptability to downstream clinical tasks than diffusion-based models. A tokenizer trained on medical images and separate semantic and structural controls enable generation across six anatomical regions in computed tomography and magnetic resonance imaging. Trained on 438,905 slices from 40 datasets, including seven internal clinical centres, MedVAR generates images 12-19 times faster than 100-step diffusion baselines. Generation quality improves with model size. Pretraining on generated images improves seven-centre hepatocellular carcinoma segmentation Dice from 0.615 to 0.632, while reconstruction using MedVAR images approaches the volumetric segmentation performance of fully sampled volumes. Membership inference reaches 3.54% sensitivity at a 1% false-positive rate, while copy detection performs near chance. These findings establish next-scale autoregression as a scalable and versatile approach to medical image generation and downstream analysis.
Figures & tables
Figure 1: A next-scale autoregressive generator across anatomical regions. a , Anatomy and modality coverage. b , Distribution of 482,815 images from 40 datasets. c , Conditioning, generation and evaluation workflow. d , Seven-centre HCC MRI cohort; labels give volume counts and red contours mark tumours. e , FID comparisons within anatomical groups and a descriptive measure of quality and efficiency normalized to depth-30 MedVAR (Methods).
Figure 2: Generation quality, efficiency and diversity across MedVAR model sizes and generator families. a , Matched examples. b , Model depths. c , FID ranks within anatomy–modality groups using MedVAR with semantic conditions alone. d , FID versus latency; MedVAR depths 16 and 20 approximately match diffusion models at 330 and 600 million parameters, respectively. Diffusion curves span 10–100 steps; paired MedVAR points use semantic conditions alone or both semantic and structural conditions. e , Fidelity across generator families. f , Condition recognizability and CAS; the dashed line marks random balanced accuracy ( 1/9 ). g , Diversity under fixed conditions. N.A., not available; error bars, 95% bootstrap CIs.
Figure 3: Medical tokenization and conditioning support explicit control. a , Natural and medical tokenizer comparison. b , Progressive conditioning. c , Conditioning ablation. d , Ablations of structural routing, the scale gate and the auxiliary structure head; MedVAR routes token maps of sizes 52 , 62 , 82 and 102 . e , Effect of the weight applied to structural conditions. f , CFG and truncation ablations. Filled and hatched bars denote generation with semantic conditions alone and with both semantic and structural conditions, respectively; arrows indicate lower is better.
Figure 4: Generated images support segmentation and reconstruction of missing slices. a , HCC segmentation and reconstruction of missing slices. b , Dice gain for each centre over pretraining on real images; points denote matched seeds. c , Real images, images generated with centre labels and unconditional MedVAR images. d , PCA and t-SNE of RadImageNet features. e , Accuracy of centre prediction and maximum mean discrepancy (MMD); the dashed line is random. f , Matched CT reconstructions (top), pairwise signed differences for Linear–U-Net, U-Net–MedVAR and MedVAR–Real on a common scale (middle), and mean CT and MRI PSNR across volumes (bottom). Linear interpolation is deterministic, the U-Net is an independent residual baseline, and circles denote the three training seeds for the system using MedVAR. g , Paired 3D Dice distributions for linear interpolation, the residual U-Net, reconstruction using MedVAR and fully sampled input; boxes show medians and interquartile ranges, whiskers extend to 1.5 times the interquartile range, and points denote volumes. Arrows indicate better performance.
Figure 5: Privacy analyses and effects of data source labels. a , White-box membership inference by ROC AUC and TPR at 1% FPR; error bars are 95% bootstrap CIs. b , ROC curves at low FPR; dashed lines mark 1% FPR and random performance. c , Nearest-neighbour copy detection for three sizes of the generated set. Diamonds and bars show mean ± s.d. across encoders; circles show individual encoders. d , Generated images and nearest training and test images with NCC, SSIM and difference maps. e , Stress test with and without labels identifying the data source under simulated contamination; positive interaction values indicate attenuation when the label is provided.
Figure 6: MedVAR architecture and semantic and structural conditioning. a , Coarse-to-fine token prediction with semantic and optional structural conditions; the auxiliary structure head is used only during training. b , Hierarchical semantic conditioning combines region and organ labels, then fuses this representation with imaging modality. A separately projected data source embedding is added at reduced weight before adaptive layer normalization. c , Structural attention routing embeds the organ mask and its boundary at each token-map scale. These features update the query and key projections of self-attention; the value projection is unchanged.