Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the background behind glass rather than the surface itself, while depth sensors such as LiDAR, time-of-flight, and RGB-D often return invalid or background measurements in transparent regions. As a result, systems that rely solely on optical sensing may misinterpret glass walls, doors, or mirrors as free space, compromising safe and reliable navigation. Existing glass segmentation approaches address this by learning visual cues such as reflections, boundaries, and semantic context from RGB images. While effective under favourable lighting and viewing conditions, these cues degrade in low-light environments, under glare, or when glass surfaces are featureless or partially occluded. In this work, we propose a multimodal framework that fuses millimetre-wave radar with RGB-D sensing for real-time transparent surface segmentation. Radar reflects strongly off glass surfaces, providing a geometric cue that remains reliable precisely where vision and depth fail. We exploit this cross-modal inconsistency to generate a radar-guided spatial prior, which is integrated into a lightweight transformer-based segmentation network, GlassFormer, via cross-modal attention. We report results on a mixed-condition test split covering all scene types and a dedicated low-light split designed to stress vision-only methods. GlassFormer achieves 0.88 mIoU on the mixed split, and 0.59 mIoU on the low light split, demonstrating substantial robustness gains over vision-only baselines while maintaining real-time performance on resource-constrained platforms.
Figures & tables
Fig. 1: RGB-D and LiDAR fail on glass, returning background depth or invalid returns through the surface, while mmWave radar reflects strongly off it. This cross-modal disagreement yields a radar-guided prior (pink) over likely transparent regions, which GlassFormer fuses with RGB to segment glass reliably, even in low light.
Fig. 2: Radar-guided mask generation pipeline. The radar produces a one-dimensional range magnitude profile from which the dominant reflection peak is extracted to estimate the radial distance dr of the nearest surface. This distance corresponds to the most prominent reflector along the radar boresight. The radar field-of-view (FoV) is then projected into the RGB image plane using camera intrinsics, defining a region of interest centered at the principal point. Within this projected FoV, the radar-derived range is compared against per-pixel depth measurements from the RGB-D sensor. Pixels with invalid depth values or depths significantly larger than the radar-confirmed surface are labeled as transparent candidates, since the depth sensor observes background geometry through the surface while the radar detects the physical interface. The resulting binary mask serves as a radar-derived transparency prior that highlights regions where optical sensing is unreliable.
Parameter
Value
Start point
120
Number of range points
200
Step length
24
Sweeps per frame
3
HWAAS
64
Receiver gain
13
TABLE I: Radar configuration for data acquisition
Fig. 3: GlassFormer architecture. A SegFormer-B2 encoder extracts a 4-scale RGB feature pyramid ( H/4 to H/32 ). At the two lowest-resolution stages, a 2-layer CNN encodes the radar mask into queries, while RGB features serve as keys and values in a RadarAttention (RA) module; a learnable spatial gate λ controls how strongly the radar prior modulates each RGB feature, letting the model down-weight radar noise. The two RA-modulated stages and the two unmodified high-resolution stages are concatenated and decoded by an MLP followed by two 1×1 convolutions to produce the segmentation mask.
Fig. 4: Qualitative comparison across scenes (top to bottom: window, occluded glass wall night outdoor, glass door, elevator glass in low light, night outdoor, mirror) against baselines.
Method
mIoU ↑
MAE ↓
F-measure ↑
BER ↓
GDNet [ 9 ]
0.5485
0.2894
0.6951
0.2618
GlassSemNet [ 3 ]
0.6336
0.1945
0.7769
–
SegFormer [ 27 ]
0.8799
0.0938
0.9361
0.0517
Radar Overlay
0.4687
0.3474
0.6382
0.3348
Ours (GlassFormer)
0.8818
0.0835
0.9372
0.0517
TABLE II: Comparison with state-of-the-art glass segmentation methods in well-lit conditions
Method
mIoU ↑
MAE ↓
F-measure ↑
BER ↓
GDNet
0.4824
0.4031
0.682
0.3981
GlassSemNet
0.4623
0.3620
0.7010
–
SegFormer
0.4912
0.3588
0.6588
0.2819
Radar Overlay
0.5407
0.3148
0.7019
0.3097
Ours (GlassFormer)
0.5904
0.2895
0.7424
0.2406
TABLE III: Comparison with state-of-the-art glass segmentation methods in low light conditions
Fig. 5: GlassFormer across viewing angles and distances. Segmentation stays reliable across the radar–RGB-D shared FoV, not just frontal viewing