Automated quality assessment of rehabilitation exercises relies heavily on accurate human pose estimation from video data. Although numerous RGB-based pose estimation methods have been proposed, the impact of camera placement on detecting clinically relevant movement errors remains insufficiently explored. To address this gap, we introduce REHAB26-ViewAngles, a dataset comprising correct and incorrect rehabilitation exercise executions captured from a wide range of camera angles. Furthermore, we propose a novel separability metric to quantify an algorithm's ability to distinguish between valid and faulty exercise repetitions. Using these tools, we analyze how various RGB-based pose-estimation strategies are suitable for exercise quality assessment under varying camera placements. In particular, we analyze single-camera 2D and 3D pose estimation and four multi-camera strategies: a combination of two orthogonal 2D views, 3D triangulation, weighted 3D fusion, and an AI-based pose-estimation transformer model specifically trained from two synchronized cameras. Our findings reveal that an optimally placed 2D camera can improve the separability by 16.9% over the commonly used 0° frontal view and frequently outperforms single-camera 3D estimation, while combining two views can further improve accuracy by up to 13.1%. These results offer practical guidance for deploying rehabilitation monitoring in both home and clinical settings.
Figures & tables
Figure 1: High-level scheme of personalized monitoring of rehabilitation exercises. The highlighted red parts are the main focus of this paper – we are especially interested in the optimal positioning of camera(s) to capture the patient’s movements.
Figure 2: The REHAB26-ViewAngles data acquisition setup. Three cameras are positioned at fixed angles ( θCam1=−90∘ , θCam2=−45∘ , θCam3=0∘ ) while the participant systematically changes their standing orientation (angle α with respect to Cam3). At each designated orientation, the participant performs a continuous set of 10 repetitions, consisting of 5 correct and 5 erroneous repetitions, each with a specific clinical error. The camera viewpoint ( VPCami ) for any given recording is mathematically defined as the sum of the participant’s rotation and the static camera position ( VPCami=α+θi ).
Exercise
Error Description
Feature Code
What Angle is Measured?
Ex1
Arm positioned too far forward
F1
Angle between the hand, shoulder, and torso
Ex1, Ex2
Shoulder elevation
F2
Angle between both shoulders, and hip
Ex1
Lateral head tilt
F3
Angle between the head, the torso, and vertical axis
Ex1, Ex2
Elbow flexion (frontal)
F4
Angle between the upper arm and forearm
Ex1, Ex2
Wrong arm positioning
F5
Angle between the hand and both shoulders
Ex2, Ex3
Lateral torso bending
F6
Angle between the torso and the lower limbs
Table 1: Overview of movement errors, their associated joint-angle measurements, and the corresponding feature codes. Some features (F2, F4-F6) are shared across multiple exercises. Features F9–F14 are defined separately for the left and right sides.
Figure 3: Representative examples from the REHAB26-ViewAngles dataset. (a) Temporal progression of Exercise 1 (Arm Abduction) recorded from a single camera. (b) Exercise 2 (Shoulder Extension) recorded simultaneously from three distinct camera perspectives. (c) Exercise 3 (Squat) captured from the same camera across three different patient orientations, with the viewpoint shifting by 10-degree increments.
−90∘
−80∘
−70∘
−60∘
−50∘
−45∘
−40∘
−35∘
−30∘
−25∘
−20∘
−15∘
−10∘
−5∘
0∘
5∘
10∘
15∘
20∘
25∘
30∘
35∘
40∘
45∘
50∘
60∘
70∘
80∘
90∘
Cam1
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
Cam2
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
Cam3
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
Table 2: Overview of all viewpoints available in REHAB26-ViewAngles dataset. A check mark indicates that the viewpoint (column header) was recorded for the given camera.
Figure 4: Comparison of a feature F2 value (angle) across Exercise1 poses for correct vs. erroneous repetitions, including the mean and 95% confidence interval of the correct repetitions.
Figure 5: Possible approaches to pose data estimation in single and multiple camera environments.
Figure 6: Illustration of feature separability scores ( Ψk ) for Exercises 1, 2, and 3 at different viewpoint angles. In each heatmap, rows correspond to camera angles and columns to individual features. The uppermost heatmaps present features computed from 2D poses. The second row of heatmaps shows features computed from 3D poses. In the last row, rotation heatmaps visualize how deviations from the optimal viewing angle affect the separability scores. Darker green indicates higher separability scores, while dark red corresponds to situations where correct and incorrect repetitions are indistinguishable.
Cams
Dims
Method
-90
-60
-40
-30
-20
-15
-10
0
10
15
20
30
40
60
90
S
2D
MPP
0.36
0.39
0.55
0.57
0.56
0.57
0.54
0.58
0.60
0.66
0.62
0.61
0.61
0.63
0.41
3D
MPP
0.31
0.44
0.49
0.46
0.51
0.57
0.50
0.47
0.54
0.55
0.57
0.59
0.47
0.51
0.42
2D
YOLO
0.21
0.38
0.55
0.60
0.53
0.57
0.53
0.56
0.61
0.63
0.64
0.62
0.60
0.59
0.41
M
2D
Ortho.
0.58
0.61
0.65
0.66
0.64
–
0.63
0.66
–
–
–
–
–
–
–
3D
Triang.
0.43
0.61
0.67
0.63
0.60
–
0.60
0.54
–
–
–
–
–
–
–
3D
Merged
0.55
0.63
0.63
0.61
0.58
–
0.62
0.51
–
–
–
–
–
–
–
Table 3: Overall separability scores across single (S) and multi-camera (M) methods and various viewpoint angles for all three exercises. For multi-camera approaches, the column header corresponds to the Cam1 viewpoint; the Cam3 viewpoint is always orthogonal (i.e., increased by 90∘ ). Higher values indicate better distinguishability between correct and incorrect movements. Bold values denote the highest separability score within each row. Underlined values highlight the absolute best score within the single-camera and multi-camera configurations, respectively. The results indicate that a single camera is often sufficient for accurate assessment, provided the patient maintains the optimal standing orientation.
Figure 7: Repetition analysis of feature F3 (Lateral head tilt) from Exercise 1 across three simultaneous viewpoints ( VPCam1=−90∘ , VPCam2=−45∘ , and VPCam3=0∘ ). Baseline values shift distinctly with the camera perspective. In the middle plot of (a), normal repetitions spread out to create a wider confidence interval, yet the erroneous repetition trajectory remains completely separate and easily detectable. Conversely, Person 2 in (b) shows a smaller separation between correct and erroneous trajectories due to a less pronounced error execution, highlighting the need for optimal camera placement. Across all setups, the frontal 0∘ viewpoint provides the most distinct signal, while a comparison between (a) and (c) demonstrates that MediaPipe’s 3D estimates introduce substantial noise and signal instability compared to 2D tracking.