Image compression, restoration, and generation all require a way to measure how different two images look to a person. Pixel error ignores how people see, while the most accurate perceptual distances are typically fitted to human judgments, tying them to a fixed data and resolution. For example, when image resolution is doubled, the correlation of DISTS with human scores on TID2013 drops from 0.815 to 0.717. We introduce the Jacobian Lens Distance (JLD), which derives its perceptual geometry from a frozen vision encoder rather than from human labels. JLD combines the locality of early patch features with the perceptual sensitivity captured by later encoder representations. Specifically, we use the encoder Jacobian to identify directions in the early feature space that most strongly affect the encoder output, producing a fixed metric tensor, E[J⊤J], which we call the Jacobian lens. The lens is fitted only once from 100 unlabeled images, taking about 35 seconds. Locally, this construction defines a pullback metric in pixel space, giving JLD a clear geometric interpretation that can be directly analyzed on real images. Across four standard perceptual databases, JLD achieves state-of-the-art performance and consistently outperforms LPIPS, DISTS, PieAPP, and DreamSim. JLD is also robust to changes in image resolution, on TID2013, its lens-term correlation remains nearly unchanged when the resolution is doubled, decreasing only from 0.850 to 0.845. We further introduce JLD-fast, which is 4× faster than LPIPS-VGG while achieving a mean correlation of 0.911. Finally, JLD naturally extends to video, reaching a correlation of 0.786 on Waterloo IVC 4K compared with 0.611 for VMAF.
Figures & tables
Figure 1: Top: block-1 patch features of the images are projected onto a fixed lens (JLD), fitted from the output’s sensitivity without labels; Bottom: five TID2013 distortions at equal PSNR (26.8 to 27.2 dB), ordered by human score, with the per-patch change, and rank and Kendall τ against the human ranks.
Figure 2: JLD on three planes spanned by two distortion directions of equal pixel energy around a reference (star), with the ellipse predicted by the pullback G (dashed). White ring: constant JLD; orange ring: constant MSE (25.1, 24.0 and 29.7 dB), whose four marked images are shown below with their JLD.
Figure 3: Three distortions of I19 (TID2013) at levels 1, 3 and 5, per patch. Block-1 displacement the lens keeps (horizontal) against the part it discards (vertical). Markers sit at each family’s RMS coordinates; their horizontal position is the lens term.
Cost per pair ↓
Spearman correlation with human ratings ↑
Method
GFLOPs
ms
TID2013 (diag.)
CSIQ (dev.)
LIVE test
KADID-10k test
Mean first four
PIPAL val
Learned deep distances, fitted to human judgements (similarity choices or quality scores)
LPIPS-VGG
240.6
10.9
0.670
0.883
0.932
0.724
0.802
0.612
DISTS
240.7
9.9
0.818
0.943
0.954
0.885
0.900
0.704
PieAPP
385.7
16.5
0.844
0.897
0.918
0.864
0.881
0.706
DreamSim
212.2
51.3
0.812
0.911
0.910
0.851
0.871
0.759
Table 1: Image distances compared on agreement with human ratings and on cost per pair. JLD uses no human label, has the best mean over the datasets of every distance. Bold marks the best distance in a column; underlining marks the second. † marks IEM results from subsets of 2,032 KADID and 250 PIPAL pairs.
Figure 4: Agreement with 699,534 human same-reference choices, per dataset and as the mean over the four datasets.
Figure 5: Equal JLD means equal damage. (a) Held-out KADID-10k distortions of one reference at equal JLD (top) and at equal PSNR (bottom), with human scores (MOS, 1 to 5). (b) Distortions of another reference in JLD order, with the per-patch JLD map of each. Selection rules in Section 5.2 .
Figure 6: Building JLD step by step on DINOv2-S block 1.
Figure 7: Resolution, cost and video. (a, b) SRCC at each short side and at native size; grey lines are the other ten methods, dashed is unweighted block-1 features. (c) Mean SRCC over the first four datasets of Table 1 against A100 time per pair, with the accuracy-cost frontier (dashed). (d) Video SRCC with 95% bootstrap intervals.
Figure 8: Held-out KADID-10k pair. White boxes locate the enlarged boat detail.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: The Jacobian lens. Left: output probes are backpropagated to the block-1 patch tokens of 100 unlabelled images; the top 64 eigenvectors of the averaged gradient outer products form the lens. Middle: the lens keeps the part of a feature change that moves the output. Right: over 153 perturbations at 40 dB PSNR, the lens displacement ranks the encoder output change far better than the raw block-1 displacement.
Figure A2: TID2013 SRCC of patch-token distances at each depth of DINOv2-S, raw or through the lens fitted there.
Figure A3: One calibration step. Probe gradients at every patch (left) are pooled into the 384×384 sensitivity matrix (middle), whose leading eigenvectors form the lens (right, first eight of 64).
Figure A4: Cumulative share of trM (violet) and of feature variance (grey).
Figure A5: Spectrum of M (top; 64 kept directions carry 76.5% of the trace; PR: participation ratio) and development SRCC against rank k .
Q in the lens term
CSIQ
KADID dev.
TID2013
KADID test
Dev. mean, with output
UkUk⊤ (flat, released)
0.956
0.876
0.874
0.870
0.932
M (Mahalanobis)
0.946
0.848
0.859
0.845
0.926
ΠkMΠk ( μ -weighted lens)
0.934
0.836
0.859
0.830
0.926
I (raw block-1 features)
0.836
0.710
0.670
0.718
0.859
Appendix
Table A1: Lens form. SRCC with MOS for the lens term under each choice of Q in ∥Q1/2δ∥ , together with the development mean (CSIQ, KADID dev.) after adding the output term.
Fit corpus
Overlap
SRCC
Δ×104
Disjoint half
Overlap
SRCC
Δ×104
10 images
0.877
0.913
−2.5
Half A, seed 0
0.963
0.913
−2.7
25 images
0.916
0.914
+3.7
Half A, seed 1
0.935
0.914
+1.6
50 images
0.963
0.913
−2.7
Half B, seed 0
0.943
0.914
+4.0
100 images
1.000
0.913
0.0
Half B, seed 1
0.951
0.913
−0.7
Appendix
Table A2: Fit stability. Subspace overlap with the 100-image lens and development mean SRCC (CSIQ, KADID dev.); Δ is relative to the 100-image fit.
Figure A6: Top: subspace overlap with the released lens. Bottom: sensitivity captured on an independent refit.
Figure A7: Local form G of JLD over TID2013 I19, one glyph per 28×28 cell (gratings of 1/8 cycle per pixel); the dashed circle is a pixel distance. Right: glyph radius against local contrast, and median radius per colour axis against spatial frequency (bands: interquartile range).
Figure A8: Selected 2AFC cases in which only JLD agrees with human scores. The green-framed image has higher MOS and receives the smaller JLD distance, while PSNR, SSIM, LPIPS-VGG, and DISTS prefer the other image.
Figure A9: Distance against MOS across distortion families and levels on the KADID-10k test split. Each point is the median over 65 references; dashed lines are linear fits. Bottom: mean MOS residual for each family. Red denotes under-penalisation by the distance and violet over-penalisation.
Figure A10: gMAD ( Ma et al., 2016 ) comparisons on the KADID-10k test split. The defender is restricted to images with nearly equal distance, while the attacker selects the pair it considers most different. The attacker wins when human scores significantly prefer its closer image (green frame).
TID2013
CSIQ
LIVE
KADID test
PIPAL pub.
PIPAL val.
Metric
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
SRCC
PLCC
PSNR
0.687
0.679
0.809
0.828
0.873
0.868
0.673
0.678
0.407
0.415
0.255
0.292
SSIM
0.627
0.686
0.837
0.815
0.910
0.915
0.621
0.630
0.504
0.509
0.363
0.426
MS-SSIM
0.786
0.834
0.913
0.899
0.951
0.948
0.826
0.823
0.562
0.593
0.491
0.566
FSIM
0.851
0.877
0.931
0.919
0.965
0.961
0.853
0.850
0.589
0.615
0.468
0.570
VSI
0.895
0.898
0.940
0.925
0.949
0.945
0.876
0.876
0.539
0.560
0.450
0.524
Appendix
Table A3: SRCC and PLCC per dataset. Violet: JLD family; grey: fitted on human judgements. Higher is better; bold and underlining mark the best and second distinct displayed values in each column.
TID2013
CSIQ
LIVE
KADID test
PIPAL pub.
PIPAL val.
Metric
KRCC
SRCC
KRCC
SRCC
KRCC
SRCC
KRCC
SRCC
KRCC
SRCC
KRCC
SRCC
PSNR
0.496
0.686
0.599
0.806
0.680
0.870
0.485
0.675
0.276
–
0.174
–
SSIM
0.455
0.628
0.632
0.837
0.731
0.910
0.449
0.623
0.349
–
0.249
–
MS-SSIM
0.605
0.784
0.738
0.912
0.803
0.949
0.635
0.826
0.397
–
0.341
–
FSIM
0.666
0.737
0.768
0.931
0.837
0.964
0.664
0.768
0.416
–
0.323
–
VSI
0.716
0.812
0.780
0.940
0.799
0.949
0.688
0.837
0.375
–
0.309
–
Appendix
Table A4: Comparing KRCC and SRCC
Variant
TID2013 (diag.)
CSIQ (dev.)
LIVE
KADID-10k test
Mean
Full block 1
0.660
0.841
0.859
0.727
0.772
Random 64
0.660
0.841
0.859
0.726
0.772
PCA 64
0.766
0.912
0.925
0.800
0.851
Lens 64
0.857
0.956
0.942
0.866
0.905
+ pre-filter
0.875
0.956
0.945
0.870
0.911
+ output term = JLD
0.877
0.971
0.964
0.892
0.926
Appendix
Table A5: Building JLD step by step on one frozen encoder (SRCC), and other distances on the same tokens. Random directions gain nothing; the directions of M give most of the gain. Bold marks the best and underlining the second distinct displayed value. † marks a score on 1,000 TID2013 pairs.
Figure A11: At equal PSNR, JLD ranks the pair as people do while DISTS and LPIPS-VGG reverse it. Each column shows a detail crop (left) and JLD’s per-token lens displacement over the whole image (right; one colour scale; the box marks the crop).
Method
256
512
Change
Deep features
Lens
0.850
0.845
− 0.005
PieAPP
0.618
0.700
+ 0.081
LPIPS-VGG
0.757
0.661
− 0.096
DISTS
0.815
0.717
− 0.099
DeepDC
0.806
0.695
− 0.111
Appendix
Table A6: TID2013 SRCC at 256 and 512 px short side. The lens remains stable while several deep feature distances degrade.
Figure A12: Which parts of JLD carry the agreement. (a) Building the distance on DINOv2-S block 1 (mean SRCC, four datasets). (b) Development SRCC against rank k . (c) Weightings of the retained directions. (d) Lens (violet) against raw features (grey) for 14 frozen encoders.
Figure A13: Mean development SRCC over CSIQ and KADID-10k development as a function of retained rank and encoder block. Bottom: block-1 lens maps at several ranks. Diagnostic fits use the preliminary combined output.
Form
TID2013 (diag.)
CSIQ (dev.)
LIVE
KADID test
KADID dev
Mean
Lens alone ( λ=0 )
0.875
0.956
0.945
0.870
0.876
0.911
Sum ( λ=0.25 )
0.877
0.969
0.959
0.885
0.887
0.923
JLD ( λ=0.5 )
0.877
0.971
0.964
0.892
0.892
0.926
Sum ( λ=0.75 )
0.875
0.970
0.966
0.895
0.895
0.927
Sum ( λ=1 )
0.872
0.968
0.967
0.897
0.896
0.926
Sum ( λ=2 )
0.861
0.959
0.965
0.897
0.894
0.920
Appendix
Table A7: Moderate output weights give similar four-dataset means. Violet: selected settings; white: controls. Bold and underlining mark the best and second displayed values.
Figure A14: Grating probes of the lens on 16 DIV2K crops. Top: peak-normalised sensitivity across spatial frequency for luminance, red–green, and blue–yellow gratings, compared with human references. Bottom left: sensitivity to oblique relative to cardinal gratings. Bottom right: local lens gain for a luminance grating (bright denotes higher sensitivity).
Figure A15: Human-study agreement. Left: one rated-image trial in which all ten participants and JLD chose A and DISTS chose B. Right: share of decided trials on which each metric agrees with the human majority (exact 95% binomial intervals; dotted: chance).
Figure A16: Approximate metamers, random perturbations, and lens-maximising perturbations at nominal 40 dB PSNR. Right: distances relative to the same-reference random perturbation; the dashed line denotes 1.
Selected block
Best block
Encoder
Block
Rank
Lens
Raw
PCA 64
Δ raw
Δ PCA
Raw
PCA 64
DINOv2-S
1
64
0.913
0.782
0.858
+0.132
+0.056
0.807
0.858
DINOv2-B
1
64
0.908
0.804
0.859
+0.104
+0.049
0.804
0.859
DINOv2-L
4
32
0.906
0.829
0.891
+0.077
+0.015
0.829
0.895
DINOv2-G
4
16
0.896
0.785
0.875
+0.112
+0.022
0.795
0.875
DINOv2-S-reg
1
64
0.902
0.818
0.871
+0.084
+0.031
0.820
0.871
Appendix
Table A8: Development-selected projections across encoder families. Violet: lens; white: controls. Bold and underlining mark the best and second displayed values in each comparison.
Figure A17: Four same-reference pairs in which humans prefer A. JLD agrees in the top two rows and reverses the preference in the bottom two. Maps share one scale within each row and show the lens term only.
Fitted to
Waterloo 4K (dev., 240)
AVT-UHD (held out, 120)
AVT
Metric
video MOS
SRCC
Within
SRCC
KRCC
PLCC
Within
Lower
encoded
JLD-video
no
0.786
0.967
0.912
0.761
0.918
0.944
0.786
0.582
VMAF
yes
0.611
0.834
0.932
0.790
0.939
0.953
0.611
0.663
MS-SSIM
no
0.694
0.967
0.884
0.718
0.888
0.945
0.694
0.561
PSNR
no
0.562
0.968
0.908
0.758
0.915
0.964
0.562
0.754
Appendix
Table A9: Video results on both datasets. Within: mean per-source SRCC. Lower: the smaller of the two dataset SRCCs. Bold marks the best value in each column.
Figure A18: Four Waterloo examples. Each row shows the reference and two distorted videos at synchronized frames; row titles report MOS and the JLD-video and VMAF choices. Coral marks disagreement with MOS and visible artifacts.
We present LatentDiff, a scalable framework for semantic dataset comparison that operates directly in the latent space of pretrained vision encoders. By combining sparse autoencoder-based divergence testing with density ratio estimation, LatentDiff identifies interpretable semantic differences between datasets at a fraction of the computational cost of caption-based alternatives. We also introduce Noisy-Diff, a benchmark capturing realistic sparse distribution shifts that cause existing methods to struggle. Experiments demonstrate that LatentDiff achieves superior accuracy while remaining robust to settings where an extremely small fraction of images (from 5% to <1% ) differ semantically.
James Flora, Kowshik Thopalli, Akshay R. Kulkarni +2
School of Electrical Engineering and Computer Science, Oregon State University, Corvallis, Oregon, USA · Lawrence Livermore National Laboratory, Livermore, California, USA · University of California San Diego, California, USA
Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human annotators' consensus. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models are at https://peterwang512.github.io/TPIPS
Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann +4
1Carnegie Mellon University · 2Adobe Research · 3UC Berkeley
We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information for downstream models. We formulate machine-oriented quality as a latent machine utility and approximate it through pairwise predictive-consistency comparisons. To this end, we construct PCMP, a dataset of PSNR-matched distortion pairs labeled by consistency votes from multiple pretrained models. We further propose ML-CLIPSim, a differentiable quality metric built on a frozen CLIP visual encoder, which aggregates intermediate patch-token similarities and global image embeddings. Experiments on machine-preference benchmarks, human-IQA datasets, and learned image compression show that ML-CLIPSim better aligns with machine-oriented preferences than conventional fidelity and perceptual metrics, while remaining competitive for human quality prediction. Used as a compression distortion term, it improves rate--task trade-offs across multiple downstream tasks.
Feng Ding, Haisheng Fu, Jie Liang +3
Simon Fraser University · University of British Columbia · Eastern Institute of Technology +2