cs.CVOct 5, 2026

JLD: Perceptual Distance Through A Jacobian Lens

Authors: Shreshth Saini, Balu Adsumilli, Alan C. Bovik

Organizations: The University of Texas at Austin, Austin, TX, USA · Google, USA · University of Colorado Boulder, Boulder, CO, USA

Abstract

Image compression, restoration, and generation all require a way to measure how different two images look to a person. Pixel error ignores how people see, while the most accurate perceptual distances are typically fitted to human judgments, tying them to a fixed data and resolution. For example, when image resolution is doubled, the correlation of DISTS with human scores on TID2013 drops from 0.815 to 0.717. We introduce the Jacobian Lens Distance (JLD), which derives its perceptual geometry from a frozen vision encoder rather than from human labels. JLD combines the locality of early patch features with the perceptual sensitivity captured by later encoder representations. Specifically, we use the encoder Jacobian to identify directions in the early feature space that most strongly affect the encoder output, producing a fixed metric tensor, E[J⊤J]E[J^\top J], which we call the Jacobian lens. The lens is fitted only once from 100 unlabeled images, taking about 35 seconds. Locally, this construction defines a pullback metric in pixel space, giving JLD a clear geometric interpretation that can be directly analyzed on real images. Across four standard perceptual databases, JLD achieves state-of-the-art performance and consistently outperforms LPIPS, DISTS, PieAPP, and DreamSim. JLD is also robust to changes in image resolution, on TID2013, its lens-term correlation remains nearly unchanged when the resolution is doubled, decreasing only from 0.850 to 0.845. We further introduce JLD-fast, which is 4×4\times faster than LPIPS-VGG while achieving a mean correlation of 0.911. Finally, JLD naturally extends to video, reaching a correlation of 0.786 on Waterloo IVC 4K compared with 0.611 for VMAF.

Figures & tables

Appendix figures & tables27 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images

    Apr 28, 2026James Flora, Kowshik Thopalli, Akshay R. Kulkarni +2Vision EncodersLatent Variable

  2. The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

    Jul 20, 2026Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann +4Similarity and MetricsRecent Vision-Language Models

  3. ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality

    May 10, 2026Feng Ding, Haisheng Fu, Jie Liang +3Image Quality AssessmentLearned Image Compression