cs.CVMar 4, 2026

Glass Segmentation with Fusion of Learned and General Visual Features

Authors: Risto OjalaTristan EllisonMo Chen

Organizations: Aalto University · Espoo, Finland · Simon Fraser University · Burnaby, Canada

Abstract

Glass surface segmentation from RGB images is a challenging task, with a number of applications in robotics and scene understanding. As glass lacks coherent visual characteristics, rich context and semantic information is crucial for accurate segmentation. Consequently, prior works on the task have explored utilization of foundation models and separate semantic backbones. This paper presents a novel dual-backbone architecture for glass segmentation, applying a frozen foundation model backbone in parallel with a learned backbone trained on task-specific segmentation data. The learned backbone enables the network to specialize to glass related visual clues, while preserving the general visual feature representations of the foundation model. The hierarchical multi-scale features acquired from the dual-backbone are compressed and decoded into segmentation masks. Benchmarking of the architecture was carried out on four commonly used glass segmentation datasets, achieving state-of-the-art results. Ablation studies highlight the performance gains of the dual-backbone design and demonstrate the generalizability of the architecture with different backbone choices. The model also has a competitive inference speed compared to the previous state-of-the-art method, and surpasses it when using a lighter backbone variant. The implementation source code and model weights are available at: https://github.com/ojalar/lgnet.

Explore similar work

Jul 27, 2026cs.CV

SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation

Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware sensors can detect glass, they lack modularity and have extensive hardware dependencies. Consequently, learning-based monocular depth estimation has emerged as a compelling alternative. However, domain-specific glass-aware monocular depth estimators struggle with unfamiliar indoor layouts; restricted by the severe scarcity of real-world glass depth annotations, they fail to generalize zero-shot to new settings. This motivates us to explore whether the extensive priors of text-to-image diffusion models can enable generalizable perception of transparent surfaces. We introduce SILICA, a unified pipeline leveraging these priors to jointly predict glass segmentation and glass-aware depth. This mutual information exchange establishes a robust spatial hierarchy, entirely eliminating the need for paired real-world glass depth annotations. Subsequently, we use the predicted segmentation mask to explicitly filter incorrect glass depth points from standard sensors, recovering accurate metric glass depth for downstream 3D mapping and autonomous collision avoidance. Supported by our novel Mirage 18k dataset, extensive experiments demonstrate that SILICA achieves remarkable zero-shot transfer across diverse, unseen environments, outperforming state-of-the-art models by almost 20% and setting a new benchmark for transparent surface perception.
Tarun R, Anuj Verma, Laksh Nanwani +2
May 12, 2024eess.IV

Explainable Convolutional Neural Networks for Retinal Fundus Classification and Cutting-Edge Segmentation Models for Retinal Blood Vessels from Fundus Images

Early detection of vision-threatening conditions such as diabetic retinopathy, glaucoma, and age-related macular degeneration depends on retinal fundus image analysis, but manual assessment is slow and expert-dependent. Automated convolutional neural networks classify fundus images accurately yet act as black boxes, and existing retinal vessel segmentation methods lose discriminative power under pathology and seldom exploit attention or transformer backbones. Using the FIVES and DRIVE fundus datasets, we develop a two-pipeline framework that pairs four-class disease classification with attention- and transformer-based vessel segmentation, organised in three stages: (1) FIVES images are augmented by rotation and horizontal and vertical flips and used to fine-tune eight ImageNet-pretrained CNNs: ResNet101, DenseNet169, Xception, InceptionV3, DenseNet121, InceptionResNetV2, ResNet50, and EfficientNetB0. (2) Five gradient-based explanation methods, Grad-CAM, Grad-CAM++, Score-CAM, Faster Score-CAM, and Layer-CAM, are computed on the final convolutional block of each classifier and compared qualitatively across architectures. (3) Ten U-Net variants are benchmarked for vessel segmentation: TransUNet (hybrid CNN--Transformer encoder) and Attention U-Net (gated skip connections), evaluated with ResNet50V2, ResNet101V2, and ResNet152V2 backbones, along with additional Attention U-Net configurations using DenseNet backbones, and the fully transformer-based Swin-UNet. ResNet101 gives the highest classification accuracy: 94.17% (F1 0.942) >> 88.33% for EfficientNetB0. For segmentation, the architecture ranking is consistent on both datasets: Attention U-Net >> TransUNet >> Swin-UNet. The strongest configuration is Attention U-Net with a ResNet101V2 backbone: FIVES IoU 0.722, Dice 0.838; DRIVE IoU 0.648, Dice 0.787, lifting DRIVE IoU 60.80 \rightarrow 64.83 over a prior custom U-Net.
Fatema Tuj Johora Faria, Mukaffi Bin Moin, Pronay Debnath +2
Aug 5, 2026cs.RO

From Transparent Labware Segmentation to Collision Avoidance: A Real-Time Edge-Aware Perception Pipeline

This paper presents an edge-aware instance segmentation framework that enables real-time robotic collision avoidance with transparent laboratory glassware using purely visual perception. Transparent vessels defy conventional segmentation due to refraction, specular reflection, and the absence of stable interior texture, yet their boundary contours remain comparatively reliable visual cues. Exploiting this observation, we augment a one-stage real-time instance segmentation backbone with a lightweight edge-detection branch, edge-guided attention fusion, and a parameter-free SimAM module, and further construct LabGlass-IS, a 3485-image, 21-category instance segmentation dataset of real laboratory glassware. The enhanced model achieves the highest Boundary F-score of 97.80 among compared methods, outperforming the YOLO-prompted FastSAM framework by 18.93 BF points. Furthermore, it maintains an inference speed of 7.1ms per frame and requires only 2.85% of the parameters of the closest accuracy competitor. Multi-view triangulation of mask centroids further provides 3D positions for conservative bounding-volume collision constraints. Real-robot trials achieve a 93.3% collision avoidance success rate, indicating the feasibility of the proposed perception-to-action pipeline for robot collision avoidance among fragile transparent objects. Our code is available at https://github.com/havishamy/TransYOLO_3D. Our video is available at https://havishamy.github.io/paper-videos/.
Shijun Ding, Chen Qian, Weiwei Shang +1