This paper addresses the challenge of automatic, frame-accurate event detection in table tennis videos. Current methods for estimating 3d ball trajectories and ball spin typically require that key events, such as ball-racket contacts, have already been identified in advance. This requirement makes it difficult to apply these methods to longer, unedited video recordings. To overcome this limitation, we propose EventNet, a two-stage pipeline to detect key events: (1) 2d keypoints are extracted of the upper-body poses for both players, table corners and ball center. A small keypoint transformer combines them into a compact representation that is robust to changes in viewpoint, lighting, and background clutter. (2) The temporal sequences of these frame-based representations are processed by a transformer encoder that predicts two time-to-event values for each frame, indicating how close the current frame is to the next and previous ball-racket contact. One novelty is a new, temporal cosine-like target signal. Furthermore, we introduce viewpoint augmentation via 3D reprojection and frame-rate augmentation to improve robustness and generalization. Our extensive ablation study gives deeper insights into the importance of various architectural and training aspects. Experimental results show that the proposed approach achieves an F1 score of 91.16% and a mean frame deviation between ground truth and predicted frame of 0.42 on the Latte-MV dataset and 73.08% / 1.16 on the challenging TTHQ dataset. Overall, our work demonstrates that 2d keypoint-based temporal modeling with our EventNet architecture is a promising and practical approach for automatic event detection in table tennis videos.
Figures & tables
Figure 1 . Overview of our EventNet: All keypoints per frame are individually condensed into a CLS token by the Keypoints Transformer , before the Time-to-Event Transformer determines the temporal distance to the previous and next event ( b(f′),f(t′) ). Enjoying the baseball game from the third-base seats. Ichiro Suzuki preparing to bat.
Figure 2 . Visualizations of our three variants of time-to-event target signals: (a) absolute linear with tmax=35 , (b) relative linear, and (c) cosine-like. Dashed lines show the target time-to-event signal: forward signal fc(t) in orange, backward signal bc(t) in blue, summed signal rc(t) in green. Solid lines show the predictions of one of our trained models.
Figure 3 . Construction of the input representation. (a) Keypoints from both players, the table, and the ball are extracted and concatenated. Here, 9 per player, 6 for the table, and 1 for the ball, resulting in 25 (x,y) keypoint coordinates per frame. (b) EventNet input by assembling the pose sequence across frames (projection each keypoint to size D ) and adding the CLS token of size D to each frame.
Figure 4 . Overview of the Keypoints Transformer architecture. It shows how a sequence of keypoints per frame is transformed into a CLS token and prepared as input for the temporal Time-to-Event Transformer.
Figure 5 . Architecture of the Time-to-Event Transformer that takes a CLS token sequence and produces time-to-event indicators f(t) and b(t) .
Figure 6 . Spatial player assignment logic using table geometry. The net line is defined by the vector between the front-mid and back-mid keypoints. Player identities are assigned by calculating the 2d cross product between this net vector and the vector pointing to each player’s bounding box center. The sign of the resulting scalar determines the "left" as Player 1 and "right" as Player 2, ensuring identity consistency despite camera perspective or player movement.
Mask all NaN Values
Mask Tokens for NaN Values
Latte-MV
TTHQ
Latte-MV
TTHQ
Configuration
F1 Score
L1 Diff
F1 Score
L1 Dff
F1 Score
L1 Diff
F1 Score
L1 Diff
NO, relative linear, ReLU
89.51 ± 0.67
0.42 ± 0.02
69.25 ± 2.75
1.14 ± 0.05
89.58 ± 1.00
0.38 ± 0.01
73.08 ± 3.92
1.16 ± 0.01
NO, relative linear, GELU
89.06 ± 1.56
0.42 ± 0.00
70.33 ± 5.44
1.13 ± 0.05
88.99 ± 2.30
0.41 ± 0.04
68.69 ± 0.63
1.18 ± 0.10
NO, relative linear, SwiGLUFFN
88.44 ± 1.32
0.43 ± 0.02
68.72 ± 4.34
1.05 ± 0.09
91.16 ± 1.28
0.42 ± 0.02
70.61 ± 2.11
1.25 ± 0.09
NO, cosine-like, ReLU
87.48 ± 1.81
0.39 ± 0.01
71.19 ± 1.14
1.16 ± 0.07
89.54 ± 0.18
0.40 ± 0.01
68.86 ± 4.35
1.28 ± 0.12
Table 1 . Test results on Latte-MV and TTHQ using our EventNet. The models are exclusively trained 3 times with Adam on Latte-MV using fps and viewpoint augmentation, a sequence length of T=29 and a maximum allowed frame distance of Δt=3 . The temporal encoding and the non-linearity used in the FFN blocks of all transformer layers are set as specified for each configuration. "L1 Diff" stands for mean absolute frame distance between predicted and ground truth event frames.
Latte-MV
TTHQ
Model
F1 Score
L1 Diff
F1 Score
L1 Dff
NO Base cos
86.0 ± 2.55
0.4 ± 0.02
46.4 ± 0.85
0.6 ± 0.06
NO EN cos
89.8 ± 1.61
0.4 ± 0.02
73.8 ± 2.33
1.3 ± 0.09
NO Base absLin
80.8 ± 1.56
0.5 ± 0.03
44.5 ± 0.28
0.5 ± 0.02
NO EN absLin
87.0 ± 2.13
0.4 ± 0.02
66.6 ± 5.01
1.1 ± 0.11
UB Base cos
88.3 ± 0.33
0.4 ± 0.08
36.1 ± 2.89
0.6 ± 0.07
Table 2 . Performance comparison of the TCN baseline (Base) with EventNet (EN, mask tokens for NaN values + fps and viewpoint augmentation) using the original absolute linear (absLin) and our cosine-like (cos) time-to-event signal (top/middle section using keypoint set NO/UB). Bottom section shows the effect if we use or one of the two data augmentation strategies.
Robotic table tennis has emerged as a compelling benchmark for real-time robotic perception due to its fast ball dynamics and stringent timing requirements. Accurate, high-frequency, and low-latency ball state estimation is critical for reliable trajectory prediction and timely control. Traditional frame-based cameras face an inherent trade-off: low frame rates leave temporal blind spots that miss fast-moving objects and high frame rates raise data and computational cost. Event cameras instead offer microsecond temporal resolution and, under sufficient illumination, remain largely free of motion blur even at high ball speeds. However, the community lacks large-scale datasets to develop and benchmark event-based perception in realistic sports scenarios. We address this gap by introducing the first large-scale event-camera dataset for table tennis, comprising over 1000 rallies from a diverse group of players ranging from amateurs to elite-level athletes. Each recording captures the event stream alongside 14 synchronized high-speed frame-based cameras at 200 FPS, which we use to produce 1 kHz pseudo ground-truth labels for ball position, velocity, and spin. Building on this dataset, we train a convolutional neural network robust to background player motion that jointly estimates the ball's position and velocity in the image-plane from events. Treating the predicted velocity as an additional measurement in the Kalman filter reduces bounce-point prediction error by 36% relative to a position-only baseline. Finally, we close the perception-action loop by integrating the event-based system with a Stäubli robotic arm, enabling the first real-time human-robot table tennis rallies driven by event-based perception.
We present TT4D, a large-scale, high-fidelity table tennis dataset. It provides 140+ hours of reconstructed singles and doubles gameplay from monocular broadcast videos, featuring multimodal annotations like high-quality camera calibrations, precise 3D ball positions, ball spin, time segmentation, and 3D human meshes over time. This rich data provides a new foundation for virtual replay, in-depth player analysis, and robot learning. The dataset's combination of scale and precision is achieved through a novel reconstruction pipeline. Prior methods first partition a game sequence into individual shot segments based on the 2D ball track, and only then attempt reconstruction. However, 2D-based time segmentation collapses under occlusion and varied camera viewpoints, preventing reliable reconstruction. We invert this paradigm by first lifting the entire unsegmented 2D ball track to 3D through a learned lifting network. This 3D trajectory then allows us to reliably perform time segmentation. The learned lifting network also infers the ball's spin, handles unreliable ball detections, and successfully reconstructs the ball trajectory in cases of high occlusion. This lift-first design is necessary, as our pipeline is the only method capable of reconstructing table tennis gameplay from general-view broadcast monocular videos. We demonstrate the dataset's fidelity through two downstream tasks: estimating the racket's pose & velocity at impact, and training a generative model of competitive rallies.
The AI CUP 2025 Precise Analysis of Table Tennis Smart Racket Data Competition introduced smart table tennis rackets that collect extensive player swing data, enabling research on table tennis big data. These data support in-depth analysis of players' return techniques and swing-force consistency, improving the accuracy of player skill assessment. This study focuses on six-axis sensor data collected by smart table tennis rackets and proposes TTNet, a novel deep learning model with multitask learning capabilities, to advance table tennis data analysis and related applications. TTNet combines convolutional neural networks (CNNs), residual networks (ResNet), and self-attention mechanisms to simultaneously predict four player attributes: gender, playing hand, years of experience, and skill level. We adopt a two-stage training strategy that incorporates data augmentation and task-specific loss functions to improve generalization on imbalanced data. Our approach achieved second place on the official competition leaderboard.
Ko-Hsun Chen, Xiang-Wei Ke, Hsien-Cheng Huang +1
Department of Computer Science and Engineering, Yuan Ze University, Taoyuan, Taiwan