Spherical Interpolation for Backward-Compatible Multimodal Representations
Organizations: DINFO (Department of Information Engineering), University of Florence, Italy · MICC (Media Integration and Communication Center) · University of Trento, Italy
Abstract
Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive at scale. Orthogonal post-hoc alignment can partially mitigate this problem by mapping new-model queries into the old-model gallery space. However, because independently trained models can differ in fine-grained representation structure, the orthogonal alignment remains approximate, leaving a residual angular discrepancy between the old-model query and the aligned new-model query. We study whether interpolation along the spherical geodesic between these two normalized query representations can improve retrieval without re-indexing the gallery. We characterize when this path contains an interior query direction closer to an idealized retrieval-optimal direction than either endpoint, and connect this characterization to Recall@ through a local margin-based certification result. Experiments across multiple benchmarks and model families show that post-alignment spherical interpolation improves over orthogonal alignment alone, recovering backward-compatibility in most evaluated settings. Consistent with our geometric characterization, per-query oracle analysis shows that retrieval-favorable interior points occur frequently in practice. Code is available at https://github.com/miccunifi/SLERP_backward_compatibility .
Figures & tables
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | R@1 |
|---|---|
| Old model | 40.62 |
| SVD | 42.89 |
| SLERP ( ) | 48.00 |
| SLERP ( ) | 48.61 |
| New model | 48.72 |
| Endpoints only (per-query oracle) | 54.61 |
| Method | I2T | T2I | Comp. | ||
| Old-model query | 53.36 | 34.26 | +1.06 | 0.00 | – |
| Alignment only | |||||
| SVD (Procrustes) | 51.72 | 33.77 | 0.00 | 45/90 | |
| Affine | 31.69 | 23.10 | 11/90 | ||
| Ridge | 33.96 | 28.89 | 14/90 | ||
| One-sided CCA | 33.96 | 28.89 | 14/90 | ||
| Encoder compute (GFLOPs) | Latency (ms) | |||||
| Serving configuration | Gallery | Query encoders | I2T | T2I | I2T | T2I |
| Old model | 8.82 | 5.96 | 19.93 | 19.95 | ||
| Full re-indexing | 162.03 | 13.30 | 25.59 | 20.40 | ||
| SLERP, one GPU | 170.85 | 19.26 | 31.04 | 25.76 | ||
| SLERP, two GPUs | 170.85 | 19.26 | 25.87 | 20.70 | ||
| SLERP, cached old query | 162.03 | 13.30 | 25.87 | 20.68 | ||