Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes substantially. We introduce Front2Back-ReID, a benchmark of 500 manually verified vehicle handovers from 20 recording sequences in South Africa. Each example asks a model to match a vehicle highlighted in a front-camera image to the same vehicle among at least three candidates in a later rear-camera image. We evaluate seven zero-shot vision-language models, four image-retrieval baselines, and 25 human participants. Models are tested using full front RGB images, cropped target vehicles, and binary silhouettes. The strongest VLM achieved 76.6 percent Rank-1 accuracy on target crops, compared with 74.0 percent for the frozen SigLIP2 baseline; this difference was not statistically clear. Human participants achieved 94.0 percent accuracy with full images and 92.2 percent with target crops. Under our evaluation setup, enabling reasoning improved accuracy across all three input conditions for every model evaluated in both modes. We also found that VLMs generally performed worse on full scenes than on target crops. These results show that general-purpose VLMs do not yet consistently outperform strong visual retrieval for front-to-rear vehicle matching, while humans remain substantially more reliable.
Figures & tables
Figure 1 : A highlighted front-view vehicle disappears through a blind interval and reappears in a closed rear gallery. The task is to select its unique match under full-RGB, crop, or silhouette front evidence.
Figure 2 : Overview of Front2Back-ReID. Curators correct and verify YOLOE-26 proposals from front and rear recordings, yielding 500 handovers (S 2 M 2 depth aids difficulty analysis only). A matcher must find the unique rear-gallery match of a highlighted front-left target after a blind interval, given full RGB, crop, or silhouette. We compare zero-shot VLMs, frozen retrieval baselines, and humans on the same decisions.
Figure 3 : Standard stereo cameras used for data collection.
Property
Value
Hand-labeled associations
500
Recordings
20
Task
Closed gallery, 1-of- Ki,Ki≥3 ,
Canonical views
Front-left, rear-left
Model conditions
3: RGB, crop, mask
Human reference conditions
2: RGB, crop
Table 1: Front2Back-ReID summary. Each record is one closed-gallery handover.
Figure 4 : Benchmark difficulty distributions: vehicle-only candidate count, approximate front-target range, normalized target scale, and occlusion. Range uses the 341 pairs with valid S 2 M 2 stereo depth; the other distributions use all 500 pairs.
Figure 5 : Representative benchmark inputs. Columns show the marked full-RGB query, the RGB target crop, the binary silhouette used in the mask condition, and the rear candidate gallery. The silhouette was supplied as black foreground on white background and contained no color, texture, lights, windows, or markings. Green match boxes and red distractor boxes are reader annotations in this figure only.
Figure 6 : Prompt specification used for all VLM evaluations. The upper and lower boxes give the decision rules and output schema common to all conditions; the middle boxes state the visual evidence available in each condition.
Model
Family
Access
Provider
Params
Open-family models
LLaVA-OneVision 0.5B
Open
Local
Local
0.5B
Llama 4 Scout
Open
API
Groq
17B/109B
Qwen 3.6 27B
Open
API
Groq
27B
Closed hosted
GPT-5.5
Closed
API
OpenAI
n/d
Table 2 : Evaluated VLMs. Experiments were run between 24 June and 5 July 2026 using the listed models. Parameter counts unavailable from providers are marked n/d.
Method
Full RGB
Target crop
Silhouette
Δctx
Retrieval controls
Random gallery
–
17.8 [14.4–21.2]
–
–
HSV histogram
–
47.6 [43.0–51.8]
–
–
DINOv2 ViT-B/14
–
49.4 [44.8–53.6]
–
–
SigLIP2 Base
–
74.0 [69.8–77.6]
–
–
Open-family VLMs
Table 3 : Rank-1 accuracy (%) without reasoning, with 95% BCa intervals. Retrieval controls use crops only. Gemini 2.5 Pro is excluded here because its reasoning cannot be disabled. Bold and underline mark the best and second-best non-human result per column. Δctx is full RGB minus target crop: every VLM except the C1-only LLaVA-OneVision loses accuracy when given the full scene, while humans gain slightly. On crops, the best VLMs perform on par with the frozen SigLIP2 encoder, and all remain far below humans.
Full RGB
Target crop
Silhouette
Method
Rank-1
Δr
Rank-1
Δr
Rank-1
Δr
Δctx
GPT-5.5
62.8 [58.2–66.8]
+4.4
76.6 [72.6–80.0]
+1.8
43.8 [39.4–48.0]
+7.0
−13.8
GPT-5.4 mini
58.8 [54.2–62.8]
+14.2
71.0 [66.6–74.6]
+12.6
40.0 [35.6–44.2]
+11.4
−12.2
Gemini 2.5 Pro
60.6 [56.0–64.6]
–
71.8 [67.4–75.4]
–
34.6 [30.4–38.6]
–
−11.2
Gemini 2.5 Flash
54.4 [49.8–58.6]
+3.8
67.2 [62.8–71.0]
+3.2
32.0 [27.8–36.0]
+6.8
−12.8
Human reference
94.0 [91.6–95.8]
–
92.2 [89.6–94.4]
–
–
–
+1.8
Table 4 : Rank-1 accuracy (%) with medium-effort reasoning. Δr is the change from the same model without reasoning (Table 3 ). Every model evaluated in both modes improves, but full RGB still trails the crop by 11–14 points. Bold and underline mark the best and second-best model.
Figure 7 : Target-crop Rank-1 accuracy without reasoning, by vehicle-only candidate count, front-target range, scale, and occlusion. The panel labeled “Rear-gallery size” uses vehicle-only count Vi , not gallery size Ki . The range panel uses the 341 pairs with valid S 2 M 2 depth, split into five equal-frequency bins; the others use all 500. Smaller targets and heavier occlusion generally lower accuracy; range and candidate-count trends are weaker.
Figure 8 : Target-crop accuracy without reasoning by weather, road context, and correct rear-target scale and occlusion. GPT-5.5 and Llama 4 Scout fall sharply on small or heavily occluded targets, while SigLIP2 stays flatter. Construction is omitted from the road-context panel (two pairs).
Figure 9 : Change in Rank-1 accuracy with medium-effort reasoning for complete matched runs. Differences and uncertainty intervals are paired by annotation; intervals crossing zero do not establish a directional gain.
Figure 10 : Human decision time by outcome. Points show medians and bars the interquartile range (IQR). All responses contribute, including those above 60 s.
Figure 11 : Five of the ten pairs missed in both human conditions. Columns show the full query, target crop, and rear gallery; green marks the match and red the distractors. Each pair received one judgment per condition.
Figure 12 : Human Rank-1 accuracy by difficulty stratum, pooled over full RGB and target crops (1,000 judgments from 25 participants). Uncertainty intervals resample responses. Accuracy remains high even for small or heavily occluded targets, where models degrade (Figs. 7 and 8 ).
Attribute
Count
Share (%)
Vehicle-only candidate count
Vi=3
45
9.0
Vi=4 – 5
163
32.6
Vi≥6
292
58.4
Front-target visibility
No occlusion
377
75.4
Table S1: Vehicle-only candidate counts and visibility (500 pairs). Vi differs from actual gallery cardinality.
Factor
Category
Count
Share (%)
Time of day
Day
488
97.6
Dawn or dusk
10
2.0
Night
2
0.4
Weather
Overcast
237
47.4
Clear
180
36.0
Sun glare
83
16.6
Table S2: Scene-context composition across the 500 handovers.
Figure S1 : Difficult handovers from Front2Back-ReID. Columns show the marked full-RGB query, target crop, and fixed rear gallery. Human participants saw one of the two front inputs and did not see the answer annotations. Green marks the match and red the distractors for the reader; detector classes are omitted and people are blurred. Rows show (a) many vehicle candidates ( Vi=14 ), (b) a distant front target ( ≈113 m approximate stereo depth), (c) front-target occlusion ( ≈61% ), and (d) boundary truncation.
Figure S2 : Human decision-time distribution by input condition, excluding 10 responses above 60 s. Violin width shows density; points are responses, horizontal marks are medians, and thick vertical segments are IQRs. Time is shown on a log scale.
Figure S3 : Human decision time by difficulty stratum after the 60 s cutoff, shown separately for full RGB and crops. Points are medians and bars are IQRs; counts pool both input conditions.
Figure S4 : Human decision time by scene context, correctness, and visibility after the 60 s cutoff. Incorrect responses take longer; weather, road context, and occlusion show smaller median differences.
Model cell
Outcome
Pick / GT
Evidence sentence
GPT-5.5 (R)
Correct
C1 / C1
C1 matches the target’s white low sedan silhouette with a dark side-window band.
Gemini 2.5 Pro (R)
Correct
C2 / C2
The target and candidate C2 are both large white tanker trucks.
GPT-5.5 (R)
Wrong
C3 / C1
C3 matches the target’s dark gray SUV body, chrome-accented grille, and headlight shape.
Gemini 2.5 Flash (R)
Wrong
C1 / C2
The target vehicle and C1 are both white vans with identical yellow and gold stripe livery on their sides.
Table S3 : Saved model evidence sentences from front-crop trials. “GT” is the correct rear-candidate alias; R denotes medium-effort reasoning.
Figure S5 : Front-crop pair missed by all 10 complete front-crop VLM cells. The correct rear candidate is C2 (green); C1 is the common selected distractor (red). Other candidates are drawn with thin outlines to preserve the small ground-truth box.
Model cell
Pick
Conf.
Evidence sentence
GPT-5.5 (R)
C1
0.86
C1 is the only candidate showing a large white-cab tractor with a long reddish-brown box trailer matching the target.
GPT-5.5 (no R)
C1
0.72
C1 shows the same large white tractor with a long reddish-brown box trailer silhouette as the front target.
GPT-5.4 mini (R)
C1
0.44
C1 is the only candidate with a long semi-trailer silhouette and white cab/trailer combination matching the target truck.
GPT-5.4 mini (no R)
C1
0.72
C1 best matches the front target’s white boxy vehicle shape with a light cab and long rectangular cargo body.
Gemini 2.5 Pro (R)
C1
0.92
The target is a white tractor pulling a white trailer, and C1 is the only candidate that is a large white truck.
Gemini 2.5 Flash (R)
C1
0.92
The target vehicle is a white semi-truck with a red shipping container, which precisely matches candidate C1.
Table S4 : Saved predictions for Fig. S5 . All runs chose C1; the correct answer is C2. R denotes medium-effort reasoning; confidence is the model’s reported value.
Figure S6 : Reasoning-disabled crop accuracy in five bins of valid approximate front-stereo depth ( n=341 ) and normalized front-target area ( n=500 ). Points are placed at each bin’s median depth or area. Depth shows no monotonic trend; SigLIP2 improves as target area increases.
Figure S7 : Durations of all 21 recorded sequences, including the one with no accepted pair (46 seconds to 17.2 minutes; median 4.7 minutes; approximately 1.9 hours in total).