The contemporary paradigm of scaling data annotation, crucial for developing machine learning solutions, is to hire ordinary human annotators and instruct them with expert-crafted guidelines to label data. This paradigm is laborious, tedious, and costly, motivating us to study an open problem, auto-annotation with expert-crafted guidelines (dubbed AutoExpert). We develop benchmarks by redesigning the evaluation protocol and re-annotating data with nuScenes and PandaSet, two 3D detection datasets for autonomous driving research that provide expert-crafted annotation guidelines. Their guidelines define 18 and 25 object classes, respectively, using nuanced language descriptions and a few visual examples. Following the guidelines that require using 3D cuboids to label LiDAR data, AutoExpert requires algorithms to learn on few-shot labeled images and texts to perform the task of 3D detection on LiDAR data. Apparently, the challenges of AutoExpert lie in the data-modality and task discrepancy. Nevertheless, public foundation models (FMs) serve as promising tools to tackle these challenges. To address AutoExpert, we adopt a conceptually simple pipeline consisting of three components: (1) 2D object detection and segmentation in RGB images, (2) lifting 2D detections into 3D using known sensor poses, and (3) 3D cuboids generation for the 2D detections. Within this pipeline, we enhance and evaluate a variety of methods such as open-vocabulary detectors, few-shot detectors, and self-supervised learned detectors. We also develop novel techniques, leading to refined components that boost 3D detection mAP from 12.1 to 25.4 on the AutoExpert-nuScenes benchmark.
Figures & tables
Figure 1: Excerpts of the authentic annotation guidelines of the nuScenes dataset [ 10 ] . (a) The guidelines instruct human annotators to label LiDAR points with 3D cuboids for specific object classes. (b) Each class is defined with a few visual examples and nuanced textual descriptions (ref. the red box ) without 3D annotations. Human can comprehend and apply these guidelines to draw 3D boxes. (c) We visualize the ground-truth or human-annotated 3D cuboids in the RGB image and the Bird’s-Eye-View (BEV) of LiDAR points.
Figure 2: To solve AutoExpert, we adopt a conceptually simple pipeline. Specifically, over the visual examples and textual descriptions that define object classes of interest, we adapt VLMs and VFMs for object detection and segmentation. The adapted FMs produce decent 2D detections on unlabeled RGB frames. With the known parameters of LiDAR and camera, we lift 2D detections to 3D with our developed VLM-Guided Multi-Hypothesis Testing (v-MHT) strategy, yielding 3D cuboids as 3D detections.
Figure 3: For each class name, we use a VLM (e.g., GPT-4o [ 2 ] and Qwen [ 4 ] ) to find a list of terms that match its description and visual examples in the annotation guidelines. We select the term or combined term that yields the best zero-shot detection performance of a foundational detector (e.g., GroundingDINO [ 33 ] ) on the valset. We construct a multimodal few-shot train-set using the selected terms and the images provided in annotation guidelines to finetune the detector, yielding notable improvements ( Table 2 ).
Figure 4: Generating 3D cuboids based on LiDAR points is challenging as points can be from occluders and backgrounds. For example, (a) LiDAR points projected on a bicycle foreground mask can be from the background scene through wheels; (b) points projected on a car mask can be from an occluding fence; (c) points projected on a car mask can be background through the windows.
Figure 5: Overview of the v-MHT strategy for 3D cuboid generation. It begins by prompting a VLM to infer the 3D information about a target 2D detection (cf. the left panel), including its 3D dimension d and orientation. As we find it challenging to directly prompt the VLM to output an orientation angle, we instruct the VLM to output the location of the object in the image and the visible parts of this object, With the output and the known camera extrinsic parameters, we estimate the object orientation θ (cf. the mid panel). Lastly, we initialize a 3D cuboid with the estimated dimension d and orientation θ and perform multi-hypothesis testing (MHT) to search for the final cuboid that best fits LiDAR points and the 2D detection box (cf. the right panel).
Figure 6: Visual results of auto3D on four testing examples of AutoExpert-nuScenes. For each example, we display 2D and 3D detections (i.e., the generated 3D cuboids) on the RGB image and the BEV of LiDAR data. Results show that auto3D decently detects objects that are in far field and small in size, which are challenging to detect [ 47 , 18 ] . Refer to Appx. § M for more visual results and the Jupyter Notebook for a demo.
Method
pub.
mAP 3D
NDS
ATE
ASE
AOE
AVE
AAE
Oyster [ 81 ]
CVPR’23
6.3
10.7
0.755
0.715
1.451
1.201
0.771
w/ frustum
ours
12.9
15.6
0.723
0.651
1.563
1.008
0.711
LISO [ 8 ]
ECCV’24
8.9
13.1
0.725
0.679
1.411
1.152
0.733
w/ frustum
ours
15.7
19.8
0.679
0.583
1.429
0.903
0.644
UNION [ 28 ]
NeurIPS’24
9.7
13.8
0.726
0.667
1.412
1.112
0.713
w/ frustum
ours
16.8
20.6
0.677
0.579
1.397
0.901
0.622
Table 1: Benchmarking on AutoExpert-nuScenes . The upper panel lists methods based on self-supervised 3D proposal detectors, which do not output class labels. We assign them with the class labeled predicted by our finetuned GroundingDINO (with refined class names). Moreover, we also empower them by providing our generated frustum (marked by “w/ frustum”), helping them achieve significant boosts. The mid panel lists methods based on open-vocabulary 3D detectors. We empower them by replacing their 2D detectors with our finetuned GroundingDINO with refined names (marked by “w/ ft-GD”). Nevertheless, our method auto3D significantly outperforms all the compared methods.
name
mAP 2D
mAP 3D
NDS
Detic
o-name
16.5
12.1
16.6
GD
o-name
16.9
16.1
21.3
ft-GD
o-name
20.0
16.6
21.2
GD
r-name
18.2
15.7
22.1
ft-GD
r-name
20.8
18.2
23.1
Table 2: Comparison of 2D detectors w.r.t 2D and 3D metrics. Detic (prompted with original name, o-name) [ 86 ] is the 2D detector used in CM3D [ 23 ] , serving as the baseline in this table. We test the performance of prompting or finetuning GroundingDINO (GD) with original class names vs. refined names (r-name) given by the GPT-4o model ( Fig. 3 ). Results show that finetuning GD (ft-GD) using r-name performs the best.
Class
10+C
6+C
2+C
C
C+2
C+6
C+10
1+C+1
3+C+3
5+C+5
bus
24.3
25.8
28.1
30.8
28.3
27.2
26.4
29.4
26.7
25.3
bicycle
22.5
25.1
28.6
30.1
32.4
29.0
26.7
30.8
29.4
28.4
emergency - vehicle
12.1
13.1
12.2
12.8
11.9
12.4
12.3
12.2
12.2
12.1
adult
34.4
43.7
56.6
59.3
60.2
46.8
36.1
61.2
56.5
49.3
child
4.4
4.9
4.5
3.4
2.8
2.6
1.9
3.5
2.9
2.7
construction - worker
13.6
16.3
22.9
25.6
28.6
24.3
20.5
27.9
25.1
22.3
Table 3: Analysis of sweep aggregation strategies on per-class 3D detection performance (mAP 3D ). We excerpt a few classes but provide all the results in Appx. Table 17 . “ P +C+ F ” denotes aggregating the past P sweeps, the current sweep C, and the future N sweeps; we drop P or F if not aggregating any past or future sweeps. In each row, we bold the highest number. Somewhat surprisingly, aggregation strategies greatly impacts performance on certain classes, e.g., for construction - worker , bicycle and traffic - cone , aggregating future 2 sweeps yields the highest performance gains.
v-MHT
SA.
S3D
track
mAP 3D
NDS
18.2
23.1
✓
21.9
25.2
✓
✓
22.8
25.9
✓
✓
✓
23.6
26.4
✓
✓
✓
✓
25.4
27.2
Table 4: Ablation study on 3D cuboid generation (§ 4.2 ). The first row shows results of CM3D [ 23 ] enhanced by using our finetuned GroundingDINO in 2D detection. “ SA. ” uses class-aware sweep aggregation ( Table 3 ); S3D incorporates 3D geometric cues to score generated 3D cuboids; “ track ” means using 3D tracks to refine scores of generated cuboids. Clearly, each component contributes with notable performance gains.
Figure 7: Failure cases. We show two failure cases (more in Appx. § M ). (1) The 2D detector produces individual detections in a crowd of bicycles whereas expert annotation requires annotating them altogether. These detections are false positives. (2) Although the 2D detector can detect cars and trucks well in the image, the occluding fences negatively affects 3D cuboid generation, yielding false positives.
Figure 8: A screenshot of how we search synonyms and object size prior for a given class name. We use both visual examples and textual description of a specific class available in annotation guidelines. Here, we use the construction-worker class as an example.
original class name
refined class name by GPT-4o
refined class name by Qwen
car
car
car
truck
truck
truck
trailer
trailer, container
trailer, container
bus
bus
bus
construction - vehicle
construction-vehicle
construction-vehicle
bicycle
bicycle bike
bicycle
Table 5: Refined class names for the AutoExpert-nuScenes defined classes by GPT-4o [ 2 ] and Qwen [ 4 ] . Notable combinatorial prompts occur in some categories, e.g., the best prompt of pushable - pullable is “pushable pullable garbage container” and “hand truck”. Quantitative comparisons are in Table 6 .
GD r-name by GPT-4o
GD r-name by Qwen
ft-GD r-name by GPT-4o
ft-GD r-name by Qwen
mAP 2D
18.2
18.0
20.8
20.7
mAP 3D
15.7
15.6
18.2
18.1
NDS
22.1
21.9
23.1
23.0
Table 6: Comparison of using different refined names for finetuning GroundingDINO on AutoExpert-nuScenes. We quantitatively compare performance of finetuned GroundingDINO using the different sets of refined names. Results are comparable to those in Table 2 . showing that both GPT-4o and Qwen refine class names in the sense of better adapting GroundingDINO to AutoExpert.
class
Detic (o-name)
GD (o-name)
ft-GD (o-name)
GD (r-name)
ft-GD (r-name)
mAP 2D
mAP 3D
mAP 2D
mAP 3D
mAP 2D
mAP 3D
mAP 2D
mAP 3D
mAP 2D
mAP 3D
car
58.3
31.9
52.6
25.1
56.4
29.1
51.3
26.1
54.2
25.6
truck
37.2
17.6
34.3
14.2
34.4
15.0
36.3
14.1
37.5
12.6
trailer
4.2
0.8
8.5
2.0
7.0
1.3
8.1
1.8
8.5
1.7
bus
59.0
6.4
59.0
5.7
60.2
8.3
59.7
5.3
59.8
6.4
construction - vehicle
11.1
14.7
9.9
9.7
4.5
12.0
10.2
8.9
11.0
9.5
Table 7: Comparisons of per-category results by using different finetuning strategies with a foundational 2D detectors. We report the results of zero-shot detector Detic [ 86 ] as a reference, which is used in [ 23 ] .
Method
pub.
mAP 3D
NDS
mATE
mASE
mAOE
mAVE
mAAE
CenterPoint [ 75 ]
CVPR’21
3.6
19.0
0.971
0.517
0.794
0.546
0.447
auto3D
ours
25.4
27.2
0.552
0.534
0.992
0.927
0.536
Table 8: Performance on AutoExpert-nuScenes of a model trained on Argoverse2. A CenterPoint model trained on the AV2 dataset exhibits severe performance degradation when directly evaluated on nuScenes. This demonstrates the challenge for LiDAR-based pretrained models to generalize across domains with different LiDAR sensors.
LiDAR Configuration
parameter/feature
nuScenes [ 10 ]
AV2 [ 70 ]
LiDAR Model
Single HDL-32E
Dual VLP-32C
Number of channels
32
32 × 2
Measurement Range
Max 100m, Effective 70m
Max 200m, Effective 100m
Vertical FOV
-30.67 ∘ to + 10.67 ∘
-25 ∘ to + 15 ∘
Horizontal FOV
360 ∘
360 ∘
Table 9: Specifics of different LiDAR sensors in nuScenes and Argoverse2 (AV2) for data collection . Clearly, the LiDAR sensors notably differ in measurement range, horizontal/vertical FOV, point cloud density, and vertical distribution patterns.
Figure 9: Example of the full prompt used for geometric reasoning via VLM. The template combines the visual input (image with a green bounding box) with highly structured textual instructions. It guides the Vision Language Model through a Chain-of-Thought reasoning process—leveraging baseline average sizes and contextual cues like road geometry—to output instance-specific 3D dimensions and view directions in a strict JSON format.
mAP 3D
NDS
mATE
mASE
mAOE
mAVE
mAAE
MHT
19.2
23.8
0.570
0.538
1.130
0.920
0.555
v-MHT
21.9
25.2
0.560
0.535
0.994
0.933
0.549
Table 10: Comparison of MHT using per-class cuboid vs. per-instance cuboid. OpenBox [ 27 ] adopts the former, inquiring GPT to output per-class cuboids prior to place them in 3D towards 3D detection. In contrast, our method v-MHT inquires GPT to understand the detected visual object in image and infer the per-instance 3D information. This notably helps our method performs better.
Rotation Step Size (in radian)
Rotation Step
π/40
π/30
π/20
π/10
π/5
mAP 3D
22.0
22.0
21.9
21.9
21.4
NDS
25.5
25.3
25.2
25.2
24.6
Translation Step Size (in meter)
Translation Step
0.3m
0.5m
0.8m
1.0m
1.5m
mAP 3D
22.1
21.9
21.3
21.0
19.1
Table 11: Sensitivity analysis of rotation and translation step size parameters. Based on bolded values, we set the corresponding translation and rotation step sizes as default, which provide good trade-off between computational cost and detection performance.
Method
2D Det. (sec.)
Segmt. (sec.)
VLM Inf. (sec.)
3D Gen. (sec.)
Total Time (sec.)
CM3D [ 23 ]
0.08
0.01
N/A
0.65
0.74
MHT (w/o VLM)
0.08
0.01
N/A
0.86
0.95
v-MHT (Ours)
0.08
0.01
0.40
0.15
0.64
Table 12: Wall-clock time comparison for processing a single LiDAR sweep. The evaluation is conducted on a node with four NVIDIA A100 GPUs. By utilizing VLM priors to constrain the geometric search space, the speedup in the MHT phase completely offsets the VLM inference overhead. As a result, our v-MHT achieves the fastest overall processing time.
Method
mAR 3D /Occlusion Level
mAP 3D /Distance (m)
60-100%
40-60%
20-40%
0-20%
0-10
10-20
20-30
0-50
CM3D [ 23 ]
33.5%
49.4%
58.1%
59.9%
26.2
23.9
13.9
18.2
v-MHT (ours)
34.1%
51.5%
59.6%
61.9%
30.2
26.2
15.6
21.9
(+0.6%)
(+2.1%)
(+1.5%)
(+2.0%)
(+4.0)
(+2.3)
(+1.7)
(+3.7)
auto3D
36.5%
54.0%
60.8%
63.1%
30.4
29.8
19.8
25.4
(+3.0%)
(+4.6%)
(+2.7%)
(+3.2%)
(+4.2)
(+5.9)
(+5.9)
(+7.2)
Table 13: Analysis of occlusion robustness (mAR 3D ) and distance performance (mAP 3D ). Our v-MHT method consistently outperforms CM3D, with particularly significant gains for heavily occluded objects (0-20% visibility) and distant targets (20-30m). The final method incorporating LiDAR aggregation, 3D cuboid scoring with geometric cues and tracking-based refinement demonstrates substantial improvements, especially for far-field and occluded objects.
Method
car
truck
trailer
bus
CV
bicycle
motorcycle
EV
SAM3D [ 77 ]
6.2
5.2
0.2
2.1
0.5
3.3
4.1
0.1
Oyster [ 81 ]
13.1
4.1
0.4
2.7
3.1
10.0
17.1
0.8
w/ frustum
20.1
9.0
1.2
4.6
6.8
21.1
35.7
1.9
LISO [ 8 ]
19.0
7.0
0.8
6.4
4.3
13.7
23.1
2.5
w/ frustum
25.0
12.8
1.4
11.7
7.9
25.1
42.3
4.5
UNION [ 28 ]
18.8
7.2
0.6
6.6
4.8
14.2
22.8
2.4
Table 14: Per-class 3D object detection performance (Part 1 of 2). Comparison of our proposed auto3D against baseline methods on standard vehicle and adult pedestrian categories. Best results are highlighted in bold . Here, CV and EV denote the construction_vehicle and emergency_vehicle classes, respectively.
Method
adult
child
PO
CW
stroller
PM
PP
debris
TC
barrier
SAM3D [ 77 ]
0.9
0.1
0.2
0.2
1.5
0.9
0.1
0.1
1.6
1.5
Oyster [ 81 ]
19.8
1.2
0.9
8.6
6.9
3.7
0.1
0.1
16.9
3.9
w/ frustum
42.0
2.5
1.6
18.4
15.3
6.7
0.1
0.1
37.1
8.1
LISO [ 8 ]
26.7
1.6
1.0
11.7
9.8
4.0
0.1
0.1
23.6
5.2
w/ frustum
47.2
2.9
1.8
21.4
17.9
7.3
0.1
0.1
43.3
9.6
UNION [ 28 ]
25.9
1.7
1.0
12.0
9.6
3.8
0.1
0.1
38.1
4.9
Table 15: Per-class 3D object detection performance (Part 2 of 2). Continuation of Table 14 , presenting the evaluation on specialized pedestrian types, personal mobility devices, and static road elements. Best results are highlighted in bold . PO , PM , PP , and TC denote police_officer , personal_mobility , pushable_pullable , and traffic_cone , respectively.
Figure 10: Architecture of the proposed PointNet [ 48 ] model Pϕ for 3D cuboid refinement . Here, B denotes the batch size, FC corresponds to a fully connected layer, and BN represents a batch normalization layer. The input to the model consists of 512 LiDAR points, each represented by a 9-dimensional feature vector (as defined in Equation 4 ). If the number of points in the frustum point cloud is fewer than 512, random oversampling is applied to reach the target count of 512 points. Conversely, if the point cloud contains more than 512 points, random downsampling is performed to reduce the number to 512. The model outputs dimensional offsets used to refine the size produced by our v-MHT 3D cuboid generation method.
center
size
orientation
score
mAP 3D
NDS
mATE
mASE
mAOE
mAVE
mAAE
25.4
27.2
0.552
0.534
1.133
0.927
0.536
✓
24.5
24.7
0.601
0.589
1.142
0.976
0.585
✓
25.4
28.7
0.552
0.386
1.113
0.927
0.536
✓
25.4
27.7
0.552
0.492
0.990
0.927
0.536
✓
✓
25.4
28.0
0.552
0.460
1.111
0.927
0.536
✓
23.8
23.8
0.613
0.600
1.241
0.982
0.614
Table 16: Analysis of learning to refine generated 3D cuboids. Suppose we manually prepare 3D cuboids on the LiDAR point clouds for the visual examples provided in the annotation guidelines. We use them to learn a model that takes as input a generated 3D cuboid and outputs refined cuboid. The refinement can translate the cuboid through re- centering , adjust cuboid size , tune the orientation , and re- score the cuboid. We tune the model over the validation set. Results show that learning such a refinement network on few-shot annotated LiDAR data is challenging, although learning to refine the size of generated cuboids on limited examples is beneficial.
Class
10+C
6+C
2+C
C
C+2
C+6
C+10
1+C+1
3+C+3
5+C+5
car
33.6
36.4
39.9
41.6
40.8
38.2
35.5
40.7
38.6
37.4
truck
20.7
21.6
22.7
22.7
22.8
22.3
21.4
22.9
22.7
22.2
trailer
1.7
1.7
1.8
1.7
1.7
1.7
1.7
1.8
1.7
1.7
bus
24.3
25.8
28.1
30.8
28.3
27.2
26.4
29.4
26.7
25.3
construction - vehicle
13.5
14.3
14.6
14.9
14.9
14.4
14.3
14.8
14.8
14.6
bicycle
22.5
25.1
28.6
30.1
32.4
29.0
26.7
30.8
29.4
28.4
Table 17: Analysis of sweep aggregation strategies on per-class 3D detection performance (mAP 3D ). “ P +C+ F ” denotes aggregating the past P sweeps, the current sweep C , and the future N sweeps; we drop P or F if not aggregating any past or future sweeps. In each row, we bold the highest number and highlight it if exceeding other numbers by 0.5 points. Somewhat surprisingly, aggregation strategies greatly impacts performance on certain classes, e.g., construction - worker , bicycle and traffic - cone , aggregating the past 2 sweeps yields remarkably better performance than other strategies.
mAP 3D
NDS
mATE
mASE
mAOE
mAVE
mAAE
baseline w/o SA
21.9
25.2
0.560
0.535
0.994
0.933
0.549
per-instance SA
22.1
25.4
0.559
0.537
0.994
0.928
0.544
per-class SA
22.8
25.9
0.556
0.536
0.993
0.919
0.538
Table 18: Results of LiDAR sweep aggregation methods . The baseline does not adopt LiDAR sweep aggregation (SA). We compare the method that aggregates LiDAR sweeps for static objects (named as per-instance SA) and our per-class SA. The per-instance SA improves over the baseline but underperforms our per-class SA. We conjecture the reasons are two-fold: (1) estimating static-vs-moving by per-instance SA is erroneous that has limited densification of LiDAR points for significantly improvement, (2) dynamic objects still have sparse points that make them difficult to detect, and (3) the follow-up step by v-MHT for 3D cuboid generation offers robustness, as it uses fixed instance-specific cuboids (to search for a 3D placement as the generated 3D cuboid) that effectively mitigate the effect caused by smears in aggregated sweeps.
Figure 11: Qualitative comparison of tracking-based score refinement strategies. Top: Our proposed method performs tracking in the 3D space (highlighted by the point cloud insets). It robustly maintains object identities (green arrows) across consecutive timestamps ( t−1 , t , t+1 ) despite significant scale changes in crowded scenes, enabling effective score boosting. Bottom: Relying solely on 2D visual tracking via SAM2 fails to connect the same individuals (red arrows). The severe 2D visual variations and occlusions lead to fragmented tracks, rendering temporal score refinement ineffective.
Figure 12: Flawed annotations in the original PandaSet and our re-annotations for the AutoExpert-PandaSet benchmark. Clearly, the original PandaSet annotations (ref. the top row) correspond to “ghost” objects that are entirely invisible. Our re-annotations address this issue (ref. the bottom row).
Class Name
GT Count
car
2,464
pedestrian
890
pylons
263
temporary_construction_barriers
179
pedestrian_with_object
122
cones
111
Table 19: Distribution of our re-annotation per category on AutoExpert-PandaSet. We sample 200 non-continuous frames and meticulously re-annotate a total of 4,695 instances across 25 diverse categories that are visible in the 2D images.
original class name
refined class name by GPT-4o
car
car
pedestrian
pedestrian
pylons
pylons
temporary_construction_barriers
plastic_pedestrian_barricade
pedestrian_with_object
pedestrian_with_open_umbrella
cones
tall_safety_cone
Table 20: Refined class names for the AutoExpert-Pandaset defined classes by GPT-4o [ 2 ] . Quantitative comparisons are in Table 21 .
mAP 2D
mAP 3D
NDS
GD o-name
17.3
10.1
17.7
ft-GD o-name
25.9
17.9
26.1
GD r-name
20.5
13.6
20.9
ft-GD r-name
27.8
18.4
27.6
Table 21: Comparison of 2D detectors w.r.t 2D and 3D metrics on AutoExpert-PandaSet. We test the performance of prompting or finetuning GroundingDINO (GD) with original class names vs. refined names (r-name) given by the GPT-4o model ( Table 20 ). Results show that finetuning GD (ft-GD) using r-name performs the best.
Method
pub.
mAP 3D
NDS
mATE
mASE
mAOE
OpenBox [ 27 ]
NeurIPS’25
10.9
15.0
0.856
0.491
1.185
w/ ft-GD
ours
15.9
18.9
0.820
0.461
1.076
AnnofreeOD [ 60 ]
ICCV’25
11.5
15.4
0.855
0.486
1.190
w/ ft-GD
ours
16.1
19.3
0.811
0.452
1.015
CM3D [ 23 ]
CoRL’24
12.3
16.0
0.831
0.522
1.182
w/ ft-GD
ours
16.6
19.7
0.803
0.455
1.104
Table 22: Benchmarking results on the AutoExpert-PandaSet test-set. Evaluated across 25 diverse categories on 200 frames. Due to the lack of temporal and attribute data, the NDS metric is adapted. Our auto3D significantly outperforms OpenBox, AnnofreeOD and CM3D.
Class Name
mAP 3D
NDS
ATE
ASE
AOE
Animals-Other
36.4
48.5
0.513
1.084
1.358
Bicycle
37.4
31.3
0.539
0.296
0.743
Bus
17.1
3.6
2.011
0.257
0.552
Car
40.8
35.6
0.846
0.058
0.610
Cones
41.4
34.3
0.345
0.163
0.900
Const_Signs
16.6
7.4
0.583
0.661
0.803
Table 23: Per-Class results by our auto3D on the AutoExpert-PandaSet test-set. The abbreviations used in the table are Const_Signs for Construction Signs; Emerg_Vehicle for Emergency Vehicle; Med_Truck for Medium-sized Truck; Mot_Scooter for Motorized Scooter; Oth_Veh for Other Vehicle; Ped_w/_Object for Pedestrian with Object; Pers_Mob_Device for Personal Mobility Device; Roll_Containers for Rolling Containers; Temp_Const_Barr for Temporary Construction Barriers; Tram/Subway for Tram or Subway.
Figure 13: More visual results of 2D detection and generated 3D cuboids (i.e., 3D detection) of auto3D on AutoExpert-nuScenes.
Figure 14: Failure cases of auto3D on AutoExpert-nuScenes.
Figure 15: Visual demonstrations available in the annotation guidelines of the AutoExpert-nuScenes benchmark . We present 6 out of the 18 categories, with 3 example images shown for each category. The green bounding boxes are 2D annotations for objects of the corresponding classes. It is worth noting the federated annotation: taking the emergency - vehicle category as an example, even if objects of the car category appear in the images, no corresponding annotations are provided.
Figure 16: Visual demonstrations available in the annotation guidelines of the AutoExpert-PandaSet benchmark . We present 6 out of the 25 categories, with 3 example images shown for each category. The green bounding boxes are 2D annotations for objects of the corresponding classes. It is worth noting the federated annotation: taking the Medium-sized_Truck category as an example, even if objects of the Car or Pedestrian categories appear in the images, annotations do not contain them.
Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved training-free and open-vocabulary. A promptable segmentation model (SAM3), queried with class names as text prompts, supplies instance masks in the vehicle's six surround-view cameras, and the masks are turned into metric 3D boxes using the geometry of the scene. The core is a controlled three-stage comparison on nuScenes in which 2D detection is held fixed and only the source of 3D geometry changes. Geometry predicted from images alone reaches 0.183 mean average precision (mAP) under the official protocol; fitting boxes from raw LiDAR points inside the same masks with training-free rules reaches 0.298 mAP / 0.348 nuScenes detection score (NDS) at zero labeling cost; borrowing supervised box geometry at inference time lifts the same detections to 0.413 mAP / 0.555 NDS, which locates the pipeline's largest deficit in measurement precision rather than 2D detection, while class confusion and confidence calibration survive that substitution. Reversing the direction, a three-state camera-witness rule built from the same masks improves a supervised LiDAR-only detector from 0.596 to 0.630 mAP, roughly half the gain of fully supervised camera fusion, with no training. A coverage analysis shows that SAM3 finds 84% of in-range objects with a correctly named mask; the classes that fail in the official metric are misnamed or geometrically unforgiving, not unseen.
Ömer Faruk Deniz, Mustafa Taha Koçyiğit
Institute of Data Science and Artificial Intelligence, Boğaziçi University, South Campus, Bebek, 34342, Istanbul, Türkiye
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D sensors, particularly LiDAR and stereo cameras. LiDAR presents unique challenges to integration within an MLLM, largely because of data sparsity and lack of a grid structure for the data. For similar reasons, fusion of camera and LiDAR data within an MLLM pipeline is also uncommon. However, most autonomous systems rely on LiDAR-based sensing, and incorporating 3D data has been proven to improve performance in traditional 3D scene perception tasks. This paper presents D3VL, a novel MLLM framework that integrates 2D and 3D time-series data in a single but simple architecture. The model aims to answer questions involving traffic scene understanding and safety. D3VL shows an 11% improvement in the KITTI Question-Answering (QA) dataset compared to baseline methods in processing 2D and 3D time-series data. This paper further introduces the Waymo QA dataset extension, which assesses models' capabilities in processing 3D and time-series data under diverse driving conditions. D3VL implementation code and WaymoQA extension can be found on our supplemental website: https://automotivesafety-lvlm.github.io
Heesang Han, A. Lynn Abbott, Abhijit Sarkar
Bradley Department of Electrical and Computer Engineering, Virginia Tech, USA · Sanghani Center for Artificial Intelligence and Data Analytics, USA · Virginia Tech Transportation Institute, USA
Reliable 3D perception of vulnerable road users (VRUs) such as cyclists and pedestrians is essential for their safety in urban traffic and a core requirement for autonomous driving (AD). Alongside advances in vehicle-based perception, research increasingly equips bicycles with sensors to study traffic from a perspective native to VRUs. Such platforms still rely on LiDAR detectors originally trained on vehicle data, yet annotated 3D data from a cyclist's perspective is scarce. How well these detectors generalise to this setting has not been evaluated. We present a 3D object detection benchmark of 1,027 annotated LiDAR keyframes (over 18,000 3D bounding boxes) from the FUSE-Bike platform in urban Munich. We evaluate four nuScenes-pre-trained detectors against 1,854 human-verified ground-truth (GT) boxes both in their original form and after finetuning on training labels produced by a VRU-dedicated auto-labelling pipeline that requires no manual annotation. The zero-shot domain gap is concentrated on the VRU classes. Finetuning recovers most of it, improving mean average precision (mAP) by up to 23.4 points with the largest gains on pedestrians and cyclists, and the adapted detectors even surpass the quality of the auto-labels they were trained on. The benchmark provides a reproducible baseline for VRU-centric 3D detection and shows that auto-labels are a viable substitute for manual annotation when adapting vehicle-trained detectors to a cyclist platform.
Mario Finkbeiner, Max A. Buettner, Kanak Mazumder +1
Intelligent Vehicles Lab (IVL), Munich University of Applied Sciences, Munich, Germany