cs.CVSep 27, 2026

When Does Geometric View Synthesis Help Wine Label Retrieval? A Public One-Shot Benchmark Across Self-Supervised and Vision-Language Backbones

Authors: Yueh-Cheng Huang

Organizations: Department of Computer Science and Information Engineering, National Dong Hwa University, Hualien 974301, Taiwan

Abstract

Geometric view synthesis can expand a single wine-label photograph into a training set, but its value with pretrained image encoders is unclear. We study this on a public WineSensed-derived benchmark of 1,000 classes, one enrollment photograph per class, and 4,295 real queries. With the earlier DINO vision transformer (ViT-S/16) recipe, geometric views raise top-1 accuracy from 34.1% to 62.6-63.7%, about three times the gain from two-dimensional (2D) augmentation. Frozen SigLIP 2-B already reaches 94.7%. A linear head over its frozen features gains 1.2-1.3 percentage points with the two geometric pipelines localized by the Segment Anything Model (SAM), while the other pipelines gain an inconclusive 0.3-0.6 points. Low-rank adaptation (LoRA) and validation-selected full fine-tuning show no clear gain within the reported confidence intervals; fixed-budget full fine-tuning loses 9-24 points. SAM localization supplies all six views for 99% of sources, compared with 43% for the edge-based front end. Recognition differences between the two cylinder constructions depend on the training recipe and are confounded by their crop and canvas conventions. Rendered-cylinder tests show different responses to source tilt, but an uncalibrated rim-ratio proxy establishes no corresponding trend in recognition on real photographs. An author-confirmed audit of 50 residual errors identifies 21 query-enrollment appearance mismatches, without establishing an irreducible error rate. These results support geometric synthesis for the tested self-supervised recipe and a smaller benefit through frozen-feature adaptation of the text-supervised encoder.

Explore similar work

CardsList
  1. Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention

    Jul 2, 2026Weichen Zhou, Yawen Zou, Chunzhi Gu +3Self-Supervised Vision TransformersVision Transformer

  2. Emergent Multi-View Geometry Through Self-Distillation

    Sep 30, 2026David Nordström, Thibaut Loiseau, Vincent Lepetit +3Multi-View3D Reconstruction

  3. DVSM: Decoder-only View Synthesis Model Done Right

    May 28, 2026Cheng Sun, Jaesung Choe, Min-Hung Chen +2Novel View Synthesis