Organizations: Department of Computer Science and Engineering, United International University · Department of Computer Science and Engineering, BRAC University
Cricket, often referred to as the "gentleman's game," adheres to a strict rule set for both batsmen and bowlers, where each delivery can significantly impact the match outcome. Detecting illegal bowling actions is crucial for maintaining fair play, yet it remains challenging for umpires to monitor in real time. Existing sensor-based solutions have limitations in live match scenarios, making real-time assessment difficult. This paper proposes a computer vision-based deep learning solution to detect illegal bowling actions in live cricket matches. To develop and evaluate our approach, we compiled a dataset of 62 videos featuring 11 male bowlers, capturing both legal and illegal bowling actions from multiple angles-front, back, and side. However, the dataset predominantly comprises right-handed bowlers with conventional actions. The proposed system identifies two key frames, the shoulder frame and the release frame from video footage of a bowler's delivery and analyzes the change in the bowling arm's angle between these frames. If the angle difference exceeds a predefined threshold (e.g., 15 degrees), the delivery is flagged as potentially illegal. We evaluated the system on a custom dataset and achieved a high true positive rate, suggesting the system's potential effectiveness in real-time match settings. However, further research is required to validate the system across diverse environmental conditions and larger datasets to ensure generalizability and robustness in various live match scenarios. To the best of our knowledge, this is the first AI-based computer vision method for detecting illegal bowling actions in cricket.
Figures & tables
Figure 1: ICC rule on illegal bowling action
Approach
Legal Bowling Detection
No ball Detection
Out Detection
Sensor used
Video Analysis
Computer Vision Techniques Applied
Pose Estimation
A Wearable Wireless Sensor based [ 22 ]
Yes
No
No
Yes
No
No
No
Inertial sensor orientation [ 23 ]
Yes
No
No
No
No
No
No
Application of Computer Vision [ 3 ]
No
Yes
No
No
Yes
Yes
No
Crick-net: A Convolutional Neural Network based Classification [ 17 ]
No
Yes
No
No
Yes
Yes
No
DRS Assisting Umpire Using Computer Vision [ 4 ]
No
Yes
Yes
No
Yes
Yes
No
Multi-person 3D Pose Estimation [ 24 ]
No
No
No
No
Yes
Yes
Yes
Table 1: Comparative analysis of different solutions
Figure 2: Workflow diagram of the whole process
Figure 3: Camera angle of bowling data collection
Description
Value
Total number of videos
62
Total number of bowlers
11
Videos for testing
12
Number of unique bowlers in test set
4
Number of videos for training
50
Table 2: Data collection description
Figure 4: Angle of different gestures with the shoulder
Figure 5: Annotated data by roboflow
Figure 6: Workflow of the decision making
Model
Accuracy (%)
mAP (%)
YOLOv5
85.9
86.5
YOLOv7
88.2
89.0
YOLOv8
89.4
90.6
YOLOv10
90.1
91.2
YOLOv11
91.0
92.0
Faster R-CNN
83.1
84.0
Table 3: Release Frame Detection Model Comparison
Figure 7: Visualization of correct and incorrect classification
Pose estimation for elbow angle
Release frame detection
Results
Accuracy
Precision
Recall
F1-Score
TPR
FPR
MCC
OpenPose
VGG-16
0.4167
0.5
0.4286
0.4615
0.4286
0.6
-0.169
VGG-19
0.3333
0.4444
0.5714
0.5
0.5714
1
-0.488
ResNet-50
0.5
0.5556
0.7143
0.625
0.7143
0.8
-0.0976
YOLOv8
0.7273
0.8333
0.7143
0.7692
0.7143
0.25
0.4485
RT-DETR
0.75
0.75
0.8571
0.8
0.8571
0.4
0.4781
Table 4: Performance comparison of different models
Cricket is one of the most celebrated sports world-wide, and technological advancement has become deeply embedded in how the modern game is analyzed and coached. Cricket shot classification and automated performance analysis add a further dimension to this trend. Traditional approaches rely on RGB video features or static images, which are sensitive to environmental variations such as camera angle, lighting, and background clutter, and often fail to capture the underlying biomechanics of batting actions. In this paper, we propose a system to improve cricket coaching that takes raw video data, extracts batsmen from video frames using YOLO, and extracts 3D pose data from video frames using MeTRAbs. The system produces sequential skeletal pose data of 30 body points and captures the biomechanical features of a batsman. As part of the system, we also propose a deep learning ensemble for shot classification of four shots: flick, pull, defense, and drive. The ensemble performed well, compared to existing classification works, achieving 97.68% accuracy. In addition, we analyzed the misclassification rates to identify cases where shots were incorrectly classified and examined their possible causes. Our proposed system allows novice players to obtain useful feedback, such as important joint angles relative to expert batsmen, which can also be useful for injury prevention. The shot classifier also helps track class-wise shots over time for further analysis. In addition to novice players, coaches can use the system for player evaluation.
Sourav Shome, M. D. Ashiquzzaman Rahad, Rameswar Debnath
Computer Science and Engineering Discipline Khulna University
Estimating the precise timing of batting impact is crucial for understanding the rapid sensorimotor control. However, this task is challenging for RGB cameras due to insufficient temporal resolution and motion blur. Similarly, Inertial Measurement Units (IMUs) are impractical for actual matches due to sensor intrusiveness and their limited temporal precision. To overcome these limitations, we propose a novel framework leveraging event-based cameras, which offer microsecond resolution and high dynamic range, to estimate impact timing based on the weighted centroid distance between the detected ball and bat. To address the domain gap between event frames and RGB images that degrades segmentation accuracy, we generate high-density event frames. We then introduce a mask refinement network that leverages these frames and bidirectional mask information, optimized using a novel loss function. Experiments on real-world datasets demonstrate that our method achieves superior accuracy under challenging conditions, including low-light environments and severe occlusions, outperforming baselines by reducing the Mean Absolute Error by approximately 63%.
Ryotaro Ishida, Wataru Ikeda, Ryosei Hara +3
Keio University · NTT Communication Science Laboratories
Temporal Action Localization (TAL) has been extensively studied in generic video understanding, while fine-grained sports scenarios, such as professional badminton, remain underexplored due to their complex and subtle spatio-temporal dynamics. In this paper, we focus on fine-grained TAL in professional badminton videos and introduce a new benchmark dataset, Fine-Badminton, which consists of 31 matches with 29 fine-grained stroke categories, covering 2104 rallies and 27597 annotated actions. To effectively capture the intricate motion patterns in such scenarios, we propose a Decoupling Spatio-Temporal Adapter (DSTA), which enables efficient modeling of spatio-temporal features within a parameter-efficient framework. Specifically, DSTA decomposes motion representation into three parallel branches, capturing temporal dynamics as well as vertical and horizontal spatial variations. The design allows the model to better distinguish subtle differences among fine-grained actions. Extensive experiments on both the Fine-Badminton dataset and the ShuttleSet benchmark demonstrate that the proposed method achieves state-of-the-art performance while introducing only a marginal increase in computational and parameter cost. These results validate the effectiveness and efficiency of the proposed approach for fine-grained temporal action localization.
Tianyu Wang, Junjie Wu, Jingquan Gao +1
School of Economics and Management, Beihang University, Beijing 100191, China · Key Laboratory of Data Intelligence and Management, Beihang University, Ministry of Industry and Information Technology, Beijing 100191, China