cs.CVSep 30, 2026

Spherical Interpolation for Backward-Compatible Multimodal Representations

Authors: Simone Ricci, Niccolò Biondi, Federico Pernici

Organizations: DINFO (Department of Information Engineering), University of Florence, Italy · MICC (Media Integration and Communication Center) · University of Trento, Italy

Abstract

Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive at scale. Orthogonal post-hoc alignment can partially mitigate this problem by mapping new-model queries into the old-model gallery space. However, because independently trained models can differ in fine-grained representation structure, the orthogonal alignment remains approximate, leaving a residual angular discrepancy between the old-model query and the aligned new-model query. We study whether interpolation along the spherical geodesic between these two normalized query representations can improve retrieval without re-indexing the gallery. We characterize when this path contains an interior query direction closer to an idealized retrieval-optimal direction than either endpoint, and connect this characterization to Recall@KK through a local margin-based certification result. Experiments across multiple benchmarks and model families show that post-alignment spherical interpolation improves over orthogonal alignment alone, recovering backward-compatibility in most evaluated settings. Consistent with our geometric characterization, per-query oracle analysis shows that retrieval-favorable interior points occur frequently in practice. Code is available at https://github.com/miccunifi/SLERP_backward_compatibility .

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Multimodal Representation Alignment for Cross-modal Information Retrieval

    Jun 10, 2025Fan Xu, Luis A. LeivaMultimodal AlignmentMultimodal Retrieval

  2. RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?

    May 22, 2026Arijit Ghosh, Aritra Bandyopadhyay, Chiranjeev Bindra +1Multimodal AlignmentMultimodal Retrieval

  3. Query-Conditioned Spherical Centroid Aggregation for Multimodal Retrieval

    Sep 14, 2026Ambuj Mehrish, Anindya Nag, Sebastiano VasconMultimodal RetrievalCenter