Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the background behind glass rather than the surface itself, while depth sensors such as LiDAR, time-of-flight, and RGB-D often return invalid or background measurements in transparent regions. As a result, systems that rely solely on optical sensing may misinterpret glass walls, doors, or mirrors as free space, compromising safe and reliable navigation. Existing glass segmentation approaches address this by learning visual cues such as reflections, boundaries, and semantic context from RGB images. While effective under favourable lighting and viewing conditions, these cues degrade in low-light environments, under glare, or when glass surfaces are featureless or partially occluded. In this work, we propose a multimodal framework that fuses millimetre-wave radar with RGB-D sensing for real-time transparent surface segmentation. Radar reflects strongly off glass surfaces, providing a geometric cue that remains reliable precisely where vision and depth fail. We exploit this cross-modal inconsistency to generate a radar-guided spatial prior, which is integrated into a lightweight transformer-based segmentation network, GlassFormer, via cross-modal attention. We report results on a mixed-condition test split covering all scene types and a dedicated low-light split designed to stress vision-only methods. GlassFormer achieves 0.88 mIoU on the mixed split, and 0.59 mIoU on the low light split, demonstrating substantial robustness gains over vision-only baselines while maintaining real-time performance on resource-constrained platforms.
Figures & tables
Fig. 1: RGB-D and LiDAR fail on glass, returning background depth or invalid returns through the surface, while mmWave radar reflects strongly off it. This cross-modal disagreement yields a radar-guided prior (pink) over likely transparent regions, which GlassFormer fuses with RGB to segment glass reliably, even in low light.
Fig. 2: Radar-guided mask generation pipeline. The radar produces a one-dimensional range magnitude profile from which the dominant reflection peak is extracted to estimate the radial distance dr of the nearest surface. This distance corresponds to the most prominent reflector along the radar boresight. The radar field-of-view (FoV) is then projected into the RGB image plane using camera intrinsics, defining a region of interest centered at the principal point. Within this projected FoV, the radar-derived range is compared against per-pixel depth measurements from the RGB-D sensor. Pixels with invalid depth values or depths significantly larger than the radar-confirmed surface are labeled as transparent candidates, since the depth sensor observes background geometry through the surface while the radar detects the physical interface. The resulting binary mask serves as a radar-derived transparency prior that highlights regions where optical sensing is unreliable.
Parameter
Value
Start point
120
Number of range points
200
Step length
24
Sweeps per frame
3
HWAAS
64
Receiver gain
13
TABLE I: Radar configuration for data acquisition
Fig. 3: GlassFormer architecture. A SegFormer-B2 encoder extracts a 4-scale RGB feature pyramid ( H/4 to H/32 ). At the two lowest-resolution stages, a 2-layer CNN encodes the radar mask into queries, while RGB features serve as keys and values in a RadarAttention (RA) module; a learnable spatial gate λ controls how strongly the radar prior modulates each RGB feature, letting the model down-weight radar noise. The two RA-modulated stages and the two unmodified high-resolution stages are concatenated and decoded by an MLP followed by two 1×1 convolutions to produce the segmentation mask.
Fig. 4: Qualitative comparison across scenes (top to bottom: window, occluded glass wall night outdoor, glass door, elevator glass in low light, night outdoor, mirror) against baselines.
Method
mIoU ↑
MAE ↓
F-measure ↑
BER ↓
GDNet [ 9 ]
0.5485
0.2894
0.6951
0.2618
GlassSemNet [ 3 ]
0.6336
0.1945
0.7769
–
SegFormer [ 27 ]
0.8799
0.0938
0.9361
0.0517
Radar Overlay
0.4687
0.3474
0.6382
0.3348
Ours (GlassFormer)
0.8818
0.0835
0.9372
0.0517
TABLE II: Comparison with state-of-the-art glass segmentation methods in well-lit conditions
Method
mIoU ↑
MAE ↓
F-measure ↑
BER ↓
GDNet
0.4824
0.4031
0.682
0.3981
GlassSemNet
0.4623
0.3620
0.7010
–
SegFormer
0.4912
0.3588
0.6588
0.2819
Radar Overlay
0.5407
0.3148
0.7019
0.3097
Ours (GlassFormer)
0.5904
0.2895
0.7424
0.2406
TABLE III: Comparison with state-of-the-art glass segmentation methods in low light conditions
Fig. 5: GlassFormer across viewing angles and distances. Segmentation stays reliable across the radar–RGB-D shared FoV, not just frontal viewing
Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware sensors can detect glass, they lack modularity and have extensive hardware dependencies. Consequently, learning-based monocular depth estimation has emerged as a compelling alternative. However, domain-specific glass-aware monocular depth estimators struggle with unfamiliar indoor layouts; restricted by the severe scarcity of real-world glass depth annotations, they fail to generalize zero-shot to new settings. This motivates us to explore whether the extensive priors of text-to-image diffusion models can enable generalizable perception of transparent surfaces. We introduce SILICA, a unified pipeline leveraging these priors to jointly predict glass segmentation and glass-aware depth. This mutual information exchange establishes a robust spatial hierarchy, entirely eliminating the need for paired real-world glass depth annotations. Subsequently, we use the predicted segmentation mask to explicitly filter incorrect glass depth points from standard sensors, recovering accurate metric glass depth for downstream 3D mapping and autonomous collision avoidance. Supported by our novel Mirage 18k dataset, extensive experiments demonstrate that SILICA achieves remarkable zero-shot transfer across diverse, unseen environments, outperforming state-of-the-art models by almost 20% and setting a new benchmark for transparent surface perception.
This paper presents an edge-aware instance segmentation framework that enables real-time robotic collision avoidance with transparent laboratory glassware using purely visual perception. Transparent vessels defy conventional segmentation due to refraction, specular reflection, and the absence of stable interior texture, yet their boundary contours remain comparatively reliable visual cues. Exploiting this observation, we augment a one-stage real-time instance segmentation backbone with a lightweight edge-detection branch, edge-guided attention fusion, and a parameter-free SimAM module, and further construct LabGlass-IS, a 3485-image, 21-category instance segmentation dataset of real laboratory glassware. The enhanced model achieves the highest Boundary F-score of 97.80 among compared methods, outperforming the YOLO-prompted FastSAM framework by 18.93 BF points. Furthermore, it maintains an inference speed of 7.1ms per frame and requires only 2.85% of the parameters of the closest accuracy competitor. Multi-view triangulation of mask centroids further provides 3D positions for conservative bounding-volume collision constraints. Real-robot trials achieve a 93.3% collision avoidance success rate, indicating the feasibility of the proposed perception-to-action pipeline for robot collision avoidance among fragile transparent objects. Our code is available at https://github.com/havishamy/TransYOLO_3D. Our video is available at https://havishamy.github.io/paper-videos/.
Glass surface segmentation from RGB images is a challenging task, with a number of applications in robotics and scene understanding. As glass lacks coherent visual characteristics, rich context and semantic information is crucial for accurate segmentation. Consequently, prior works on the task have explored utilization of foundation models and separate semantic backbones. This paper presents a novel dual-backbone architecture for glass segmentation, applying a frozen foundation model backbone in parallel with a learned backbone trained on task-specific segmentation data. The learned backbone enables the network to specialize to glass related visual clues, while preserving the general visual feature representations of the foundation model. The hierarchical multi-scale features acquired from the dual-backbone are compressed and decoded into segmentation masks. Benchmarking of the architecture was carried out on four commonly used glass segmentation datasets, achieving state-of-the-art results. Ablation studies highlight the performance gains of the dual-backbone design and demonstrate the generalizability of the architecture with different backbone choices. The model also has a competitive inference speed compared to the previous state-of-the-art method, and surpasses it when using a lighter backbone variant. The implementation source code and model weights are available at: https://github.com/ojalar/lgnet.
Risto Ojala, Tristan Ellison, Mo Chen
Aalto University · Espoo, Finland · Simon Fraser University +1