cs.CVDec 27, 2025

Comparing Object Detection Models for Electrical Substation Component Mapping

Authors: Namish Bansal, Haley Mody, Dennies Kiprono Bor, Dante Groccia, Edward J. Oughton

Organizations: George Mason University

Abstract

Electrical substations are a significant component of an electrical grid. Indeed, the assets at these substations (e.g., transformers) are vulnerable to hazards such as hurricanes, flooding, earthquakes, and geomagnetically induced currents (GICs). Because failures can have significant economic and public safety implications, identifying key substation components is essential for quantifying vulnerability. Unfortunately, traditional manual mapping of substation infrastructure is time-consuming and labor-intensive. Therefore, an autonomous solution utilizing computer vision models is preferable, as it offers greater convenience and efficiency. In this study, we train and compare 16 models on a manually labeled dataset of US substation images. These models include 12 You Only Look Once (YOLO) models, 2 Roboflow Detection Transformer (RF-DETR) models, and 2 Cascade R-CNN models. RF-DETR-large achieved the highest overall detection performance with mAP@50 and mAP@50:95 scores of 0.881 and 0.632, respectively. Across all models, alternate energy systems were detected most accurately, while transformers and reactors were more difficult to identify due to their smaller size and greater visual variability. Applying our best-performing model to nationwide imagery yielded approximately 22,591 component detections across 11,083 unique substations within the United States. These detections were broken down by state and Federal Energy Regulatory Commission (FERC) regions, with Florida (2,478 detections) and Midcontinent Independent System Operator (MISO; 4,329 detections) having the largest number of detections in their respective categories.

Figures & tables

Explore similar work

Aug 4, 2026cs.CV

Advancing Utility Pole and Sign Detection Through Deep Learning

Utility poles are an essential part of the infrastructure used to support power distribution systems and other critical public services. Their regular inspection is crucial to ensure the stability and safety of the electrical grid. A deep learning framework is presented for the automated detection, segmentation and lean angle estimation of wooden utility poles, and classification of attached electrical warning signs, using ground-level imagery. The system is trained on a custom dataset of 4,570 annotated images extracted from Google Street View, featuring challenging real-world scenes with visually ambiguous wooden poles lacking distinctive features. The proposed model is based on the Detection Transformer (DETR), suitably modified and trained on the custom dataset. The model outperforms standard object detectors (RetinaNet, Faster R-CNN, YOLOv3-Tiny), achieving a mean average precision of 90.43% for pole detection and 88.26% for sign detection. Extending this model with a segmentation head enables per-instance mask generation, which is then used to estimate pole lean angle. The model accurately estimates lean for 1,367 out of 1,433 test-set poles, with a mean absolute error of 1.01 degrees. Moreover, the custom dataset created in this work is also made publicly available to be used as a benchmark.
Aug 11, 2026cs.CV

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. However, many existing approaches rely on controlled datasets that do not adequately represent realworld farming conditions, particularly in underrepresented regions such as Africa. This study presents a comparative evaluation of six object detection models YOLOv5, YOLOv8, YOLO11, YOLO26, Faster R-CNN, and RT-DETR using a real-world dataset, AgriAISeg 1 , collected manually from Nigerian farms. AgriAISeg comprises 3,382 images of sesame, cabbage, and tomato crops captured under varying environmental conditions, including changes in illumination, occlusion, and viewing perspectives. Models were trained, and performance was assessed using precision, recall, [email protected], and [email protected]:0.95. The results show that RT-DETR achieved the highest overall performance with a precision of 0.768 and [email protected]:0.95 of 0.624, while YOLOv8 and YOLO11 also demonstrated strong and consistent performance. In contrast, Faster R-CNN recorded significantly lower accuracy, with an overall [email protected] of 0.466, indicating reduced effectiveness under complex field conditions. In addition, YOLO-based models exhibited superior training efficiency compared to Faster R-CNN.These findings demonstrate that modern one-stage and transformer-based detectors provide more reliable and efficient solutions for plant detection in realworld agricultural environments.
May 13, 2026cs.CV

Pattern-Enhanced RT-DETR for Multi-Class Battery Detection

Accurate and efficient battery detection is increasingly important for applications in electronic waste recycling, industrial quality control, and automated sorting systems. In this paper, we present both a comprehensive benchmark and a novel method for multi-class battery detection. We systematically compare three CNN-based detectors (YOLOv8n, YOLOv8s, YOLO11n) and two transformer-based detectors (RT-DETR-L, RT-DETR-X) on a publicly available dataset of approximately 8,591 annotated images under identical experimental conditions, and further propose PaQ-RT-DETR, which introduces pattern-based dynamic query generation into RT-DETR to alleviate query activation imbalance with negligible computational overhead. Among baselines, YOLO11n achieves the best CNN-based accuracy (mAP@50: 0.779) at only 2.6M parameters, while YOLOv8n delivers the fastest inference at ~1,667 FPS. PaQ-RT-DETR-X achieves the highest overall mAP@50 of 0.782, surpassing RT-DETR-X by +2.8% with consistent per-class gains across all six battery categories including the data-scarce Bike Battery class. Our findings provide practical guidance for selecting object detection models in battery-related industrial applications.