ProtoSeam: Lifting Classifier Training with Latent Gaussian Mixture Models
Organizations: Department of Mathematics, Otto von Guericke University (Magdeburg, Germany) · Max Planck Institute for Dynamics of Complex Technical Systems (Magdeburg, Germany)
Abstract
We propose a lifted reformulation of supervised classification that improves the final accuracy of standard classifiers without changing the architecture at inference time. A network is split at a single semantic interface and one learnable prototype per class is inserted there. Training combines a quadratic consensus penalty that pulls toward the prototype of its class with a classification loss of evaluated on samples drawn around the prototypes, whereat no gradient crosses the interface. At inference the prototypes are discarded and the unmodified network is used. Across CIFAR-10, CIFAR-100, and TinyImageNet with ResNet and vision transformer backbones, lifted training improves test accuracy by up to five percentage points over variants without lifting under a shared tuning protocol. Moreover, we provide theoretical justification of those results.
Figures & tables
| Dataset | Model | Variant | seed 42 | seed 43 | seed 44 | Mean | Spread |
|---|---|---|---|---|---|---|---|
| CIFAR-10 | ResNet-8 | Baseline | |||||
| Unlifted | |||||||
| ProtoSeam | |||||||
| ViT-S | Baseline | ||||||
| Unlifted | |||||||
| ProtoSeam |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Model | Variant | ||
|---|---|---|---|---|
| CIFAR10 | ResNet8 | Baseline | ||
| ProtoSeam | ||||
| ViT | Baseline | |||
| ProtoSeam | ||||
| CIFAR100 | ResNet8 | Baseline | ||
| ProtoSeam |