Model inversion is a widely adopted technique in data-free learning that reconstructs synthetic inputs from a pretrained model through iterative optimization, without access to original training data. Unfortunately, its application to state-of-the-art Vision Transformers (ViTs) poses a major computational challenge, due to their expensive self-attention mechanisms. To address this, Sparse Model Inversion (SMI) was proposed to improve efficiency by pruning and discarding seemingly unimportant patches, which were even claimed to be obstacles to knowledge transfer. However, our empirical findings suggest the opposite: even randomly selected patches can eventually acquire transferable knowledge through continued inversion. This reveals that discarding any prematurely inverted patches is inefficient, as it suppresses the extraction of class-agnostic features essential for knowledge transfer, along with class-specific features. In this paper, we propose Patch Rebirth Inversion (PRI), a novel approach that incrementally detaches the most important patches during the inversion process to construct sparse synthetic images, while allowing the remaining patches to continue evolving for future selection. This progressive strategy not only improves efficiency, but also encourages initially less informative patches to gradually accumulate more class-relevant knowledge, a phenomenon we refer to as the Re-Birth effect, thereby effectively balancing class-agnostic and class-specific knowledge. Experimental results show that PRI achieves up to 10x faster inversion than standard Dense Model Inversion (DMI) and 2x faster than SMI, while consistently outperforming SMI in accuracy and matching the performance of DMI.
Figures & tables
Figure 1: Comparison of DMI [ 41 ] , SMI [ 19 ] , and PRI (ours) on CIFAR-100. (a) Student accuracy is distilled from DeiT-Tiny/Small/Base (T/S/B) teachers. GPU time is measured in log10 minutes for inverting 128 samples per batch. Red dashed lines represent teacher accuracy. (b) Visualization illustrating the visual fidelity differences among DMI [ 41 ] , SMI [ 19 ] , and our proposed PRI.
Table 1: Knowledge distillation performance under different patch selection strategies in SMI: high-attention, low-attention, random, and fixed-region (top), where DeiT-Base is fine-tuned on 32 inverted CIFAR-10 images with 76% sparsity for 120 epochs.
Figure 3: Overview of patch rebirth inversion. At each iteration ti , we store blue framed important patches and mask them out (black) while the remaining red framed patches continue inversion, progressively embedding class-specific features. All stored sparse view compose the final synthesized dataset, which is used for data-free downstream tasks.
Figure 4: Confusion matrices of student models trained exclusively on inverted images from a single class, “airplane”, in CIFAR-10, using different inversion methods. The architecture of both teacher and student is DeiT-Base. While students trained with DMI and SMI fail to generalize beyond the target class, the student trained with PRI-inverted images exhibits broad generalization across all classes.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: t-SNE visualization of feature embeddings from DMI, SMI, and multiple detachment points of PRI on CIFAR-100. All embeddings are extracted using the pretrained DeiT-Base.
Figure A2: Learning curves on CIFAR-100 using a DeiT-Base teacher and a DeiT-Base student. (a) Comparison of learning curves with inversion time among three inversion methods (PRI with 75% sparsity, SMI with 76% sparsity and DMI). (b) Learning curves using samples generated at different detachment points from t1 to t4 .
Figure A3: Re-birth visualizations. For each class, the left image shows the initially regarded as unimportant patches, while the right image shows the same patches after further inversion.