What Do Scan-Derived Class Prototypes Add? Disentangling Supervision, Prototype Content and Query Protocol in Recognition over Frozen Foundation Features
Organizations: George W. Woodruff School of Mechanical Engineering, Georgia Institute of Technology, Atlanta, 30332, GA, USA
Abstract
A scan supplies labeled images and a geometric reference. We separate their contributions in a recognizer whose scan-derived prototype matrix acts as a supervised head's fixed output layer. On T-LESS, HOPE and 18 self-collected industrial parts, we test real, random and exactly permuted prototypes, matched geometry-free classifiers, stronger appearance rules and paired background protocols. Across DINOv2-giant and MetaCLIP-H with real-background queries, the largest fused-accuracy advantage of the real prototypes over either control is one percentage point; larger differences favor controls, by up to 2.8 points in arm means. On HOPE with DINOv2-giant the head alone is 2.8 points above exact permutations (95% interval: 0.8-4.7); this advantage does not reach fusion and is not observed on MetaCLIP-H. On DINOv2-giant, matched logistic regression comes within 0.5 points of fusion on T-LESS and exceeds it on HOPE and the self-collected parts. Against white cutouts, real HOPE query backgrounds lower image-prototype accuracy by 43 points on DINOv2-giant and 13 on MetaCLIP-H. The audit separates prototype content, label supervision and query protocol.
Figures & tables
| Method | Query and reference | Decision path / object-specific training |
|---|---|---|
| CNOS, SAM-6D | RGB; CAD-rendered templates. SAM-6D also uses depth for pose. | Frozen template matching; none. |
| FoundPose | RGB; CAD RGB-D templates. | Template retrieval, 2D–3D correspondences and PnP; none. |
| GFreeDet | RGB; Gaussian-Splatting-rendered templates. | Frozen template matching; none. |
| FreeZe | RGB-D; model point cloud with frozen geometric and visual features. | Registration of query and model clouds; none. |
| ULIP, OpenShape | Point cloud or image; joint image–text–point-cloud embedding. | Retrieval in the joint embedding; none per object. |
| Audited system | RGB; per-class -D scan descriptors ( ). | is a fixed classifier output layer; head learns from labeled onboarding frames. |
| Dataset (classes) | Onboarding | Per class | Queries |
|---|---|---|---|
| T-LESS ( ) | |||
| HOPE ( ) | |||
| FGIP ( ) |
| Coverage | Views | Minutes | GB |
|---|---|---|---|
| One equatorial ring | |||
| Three rings | |||
| Five rings (main) |
| Dataset | Output | Real | Random | Permuted |
|---|---|---|---|---|
| T-LESS | Image-only | |||
| head | ||||
| fused | ||||
| best_val | ||||
| HOPE | Image-only | |||
| head | ||||
| Dataset | Level | Real random [95%] | Real permuted [95%] |
|---|---|---|---|
| T-LESS | head | [ , ] | [ , ] |
| fused | [ , ] | [ , ] | |
| HOPE | head | [ , ] | [ , ] |
| fused | [ , ] | [ , ] | |
| FGIP | head | [ , ] | [ , ] |
| fused | [ , ] | [ , ] |
| Queries | Training / classifier | T-LESS | HOPE | FGIP |
|---|---|---|---|---|
| Real | Matched LR | |||
| Matched softmax | ||||
| Full LR | ||||
| Full softmax | ||||
| Geometry head | ||||
| Fused |
| T-LESS | ||
|---|---|---|
| Rule | Alone | head |
| Mean prototype | ||
| Best template | ||
| Mean of best five | ||
| -means ( centers) | ||
| Soft vote, all templates | ||
| Alone | real head | random | |||
| HOPE rule | White | Real | White | Real | Real |
| Mean prototype | |||||
| MetaCLIP mean | — | — | — | ||
| Best template | |||||
| Mean of best five | |||||
| -means ( ) | |||||
| Dataset | Level | Real | Random | Permuted |
|---|---|---|---|---|
| T-LESS | image-only | / | ||
| head | ||||
| fused | ||||
| best_val | ||||
| HOPE | image-only | / | ||
| head | ||||
| Real minus, over runs ( ) | Real minus, seed | ||||
|---|---|---|---|---|---|
| Dataset | Level | random | permuted | random | permuted |
| T-LESS | head | ( ) | ( ) | ||
| fused | ( ) | ( ) | |||
| HOPE | head | ( ) | ( ) | ||
| fused | ( ) | ( ) | |||
| FGIP | head | ( ) | ( ) | ||
| Reference / arm | T-LESS | HOPE | FGIP |
|---|---|---|---|
| MetaCLIP-H, matched optimization subset | |||
| Logistic regression | |||
| Softmax head | |||
| Geometry head | |||
| DINOv2-giant: head / fused | |||
| Real ( seeds) | / | / | / |
| Assignment (runs) | Test | Train |
|---|---|---|
| Semantic ( ) | ||
| Permuted ( ) |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Encoder | Image-only | Geometry head | Fused |
|---|---|---|---|
| SigLIP2-g † | |||
| MetaCLIP-H | |||
| SigLIP-L | |||
| DINOv2-giant | |||
| SigLIP2-so400m | |||
| OpenCLIP-H |
| Arm | Pose-free | Pose-assisted |
|---|---|---|
| RGB patch BoW | — | |
| CAD surface + query depth | ||
| 3DGS scan + query depth | ||
| Onboarding depth + query depth | ||
| CAD box + query depth | ||
| Pair details: raw RGB (left); assisted CAD / 3DGS (right) | ||