cs.CVSep 27, 2025

Patch Rebirth: Fast and Transferable Model Inversion of Vision Transformers

Authors: Seongsoo Heo, Dong-Wan Choi

Organizations: Inha University

Abstract

Model inversion is a widely adopted technique in data-free learning that reconstructs synthetic inputs from a pretrained model through iterative optimization, without access to original training data. Unfortunately, its application to state-of-the-art Vision Transformers (ViTs) poses a major computational challenge, due to their expensive self-attention mechanisms. To address this, Sparse Model Inversion (SMI) was proposed to improve efficiency by pruning and discarding seemingly unimportant patches, which were even claimed to be obstacles to knowledge transfer. However, our empirical findings suggest the opposite: even randomly selected patches can eventually acquire transferable knowledge through continued inversion. This reveals that discarding any prematurely inverted patches is inefficient, as it suppresses the extraction of class-agnostic features essential for knowledge transfer, along with class-specific features. In this paper, we propose Patch Rebirth Inversion (PRI), a novel approach that incrementally detaches the most important patches during the inversion process to construct sparse synthetic images, while allowing the remaining patches to continue evolving for future selection. This progressive strategy not only improves efficiency, but also encourages initially less informative patches to gradually accumulate more class-relevant knowledge, a phenomenon we refer to as the Re-Birth effect, thereby effectively balancing class-agnostic and class-specific knowledge. Experimental results show that PRI achieves up to 10x faster inversion than standard Dense Model Inversion (DMI) and 2x faster than SMI, while consistently outperforming SMI in accuracy and matching the performance of DMI.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers

    Jul 1, 2026Aravind Pradeep, Samira Nazari, Mahdi Taheri +1Self-Supervised Vision TransformersVision Transformer

  2. Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling

    Sep 1, 2026Stefano Leggio, Giulio Rossolini, Alessandro BiondiVision TransformerToken Embeddings

  3. Attention Transfer Is Not Universally Effective for Vision Transformers

    May 8, 2026Huaiyuan Qin, Muli Yang, Gabriel James Goenawan +4Self-Supervised Vision TransformersVision Transformer