Recent 4D LiDAR language models aim to reason about objects and their evolving spatial relationships. Yet, in our evaluation, always selecting the same option nearly matches the multiple-choice accuracy of two B4DL-derived configurations. We introduce LiDAR-Hallu, a geometry-referenced benchmark and diagnostic protocol with 10,000 questions across 150 nuScenes scenes. It covers object existence, ego-relative position, distance ordering, relative motion, and temporal localization, with explicit rules for selecting objects, comparing times, and determining reference answers. Our protocol combines fixed-answer and candidate-content controls, cross-scene pairs with identical prompts but opposite reference answers, and relation-specific recall. Analysis of 100,000 recorded responses reveals failures hidden by aggregate accuracy. Candidate duration alone makes temporal answers predictable without observing LiDAR. On paired questions, the models frequently give the same answer to scenes requiring opposite answers. Relation-specific analysis further shows that both configurations miss every positive lateral-motion case across all tested conditions. Temporal-shuffle contrastive decoding provides little net improvement, as repairs are largely offset by new errors and the main failures persist. These results show that evaluating spatio-temporal reasoning requires testing whether models distinguish the queried physical relationships, rather than relying on individual-answer accuracy alone. The source code, checkpoints, and data are released at https://github.com/Awesome4D/4DMLLM_Hallucination_Bench.
Figures & tables
Fig. 2: Question and reference construction. (a) Tracking the same bus gives its ego-relative lateral change. (b) A source event interval and three distractors form a temporal-localization question.
Dataset characteristics
Diagnostic evaluation
Dataset
Input
LiDAR- specific
360∘ coverage
Multi- frame
Sequence
Sequence annotation
Answer-bias controls
Exact-prompt pairs
Relation recall
DriveLM [ 2 ]
Camera
×
✓
✓
×
×
×
×
×
LingoQA [ 10 ]
Camera
×
×
✓
✓
✓
✓
×
×
DriveGPT4 [ 11 ]
Camera
×
×
✓
✓
✓
×
×
×
nuScenes-QA [ 6 ]
Camera + LiDAR
✓
✓
✓
✓
×
✓
×
×
LiDARLLM [ 4 ]
LiDAR
✓
✓
×
×
×
×
×
×
TABLE I: Driving-scene language benchmarks: sensor input, temporal coverage, and diagnostic evaluation. Dataset characteristics follow Table 1 of B4DL [ 5 ] . The diagnostic columns indicate reported evaluations.
Family
Binary
MCQ
Reference
Existence
1,000
1,000
Class membership
Position
1,000
1,000
Ego-frame sector
Distance
0
2,000
Nearest class
Motion
2,000
0
Same-instance deltas
Temporal
0
2,000
Source interval
Total
4,000
6,000
150 scenes
TABLE II: Benchmark composition and reference answers. MCQ denotes a four-choice question.
Fig. 4: Evaluated model architecture. Projected LiDAR-CLIP features and question tokens form the input to Vicuna-1.5-7B. B4DL-Ego also includes ego-motion text.
Task
Fixed
B4DL
B4DL-Ego
Content
Existence (B)
50.00
66.30
67.20
72.10
Existence (M)
25.00
26.70
26.30
66.10
Position (B)
50.00
45.70
44.50
54.90
Position (M)
25.50
26.10
27.00
29.10
Motion (B)
50.00
52.40
50.00
59.05
Distance (M)
24.30
24.30
24.40
52.70
TABLE III: Accuracy (%) by task and answer format. B: binary; M: four-choice. Fixed answers are No for B and A for M. Content controls use statistics from other scenes.
Condition
Item acc. (%)
Both correct
Joint (%)
Prompt-only null
50.00
0/348
0.00
Baseline
53.30
43/348
12.36
Anti-position
51.87
39/348
11.21
TSCD α=0.5
52.16
42/348
12.07
TSCD α=1
52.30
42/348
12.07
TSCD α=1.5
52.01
40/348
11.49
TABLE IV: B4DL on 348 exact-prompt pairs with opposite reference answers. The prompt-only row gives the deterministic reference.
Fig. 6: Qualitative results on LiDAR-Hallu. Examples cover distance, position, existence, temporal localization, and paired motion questions. Predictions compare default decoding with TSCD.
Positive recall
Accuracy
Relation
Yes / No
B4DL
B4DL-Ego
B4DL
B4DL-Ego
Closer
206 / 209
4.37
7.77
45.78
44.58
Farther
347 / 141
30.26
7.20
43.44
33.81
Left
145 / 199
0.00
0.00
57.85
57.85
Right
157 / 221
0.00
0.00
58.47
58.47
Similar
145 / 230
55.86
61.38
60.27
61.33
TABLE V: Motion labels, positive recall, and accuracy (%) under default decoding.
TPR ↑
FPR ↓
Family
B4DL
B4DL-Ego
B4DL
B4DL-Ego
Existence
49.80
40.60
17.20
6.20
Position
69.80
63.80
78.40
74.80
Motion
19.50
13.00
14.70
13.00
TABLE VI: Positive recall (TPR) and false-positive rate (FPR), in %, under default decoding.
Model
Intervention
R
D
ΔA
95% interval
B4DL
Prompt
58
65
-0.07
[-0.27, +0.13]
B4DL
α=0.5
117
114
+0.03
[-0.25, +0.31]
B4DL
α=1
119
115
+0.04
[-0.25, +0.34]
B4DL
α=1.5
119
116
+0.03
[-0.27, +0.33]
B4DL-Ego
Prompt
50
117
-0.67
[-0.93, -0.41]
B4DL-Ego
α=0.5
161
146
+0.15
[-0.22, +0.52]
TABLE VII: Paired intervention results. R/D : repairs/regressions. Accuracy changes and 95% scene-bootstrap intervals are in percentage points. Prompt denotes anti-position prompting.
Department of Automotive Engineering (Automotive-Computer Convergence), Hanyang University, Seoul, South Korea. · Department of Automotive Engineering, Hanyang University, Seoul, South Korea.