We present the first motion generation system for playtesting virtual reality (VR) games. Our player model generates VR headset and handheld controller movements from in-game object arrangements, guided by style exemplars and aligned to maximize simulated gameplay score. We train on the large BOXRR-23 dataset and apply our framework on the popular VR game Beat Saber. The resulting model Robo-Saber produces skilled gameplay and captures diverse player behaviors, mirroring the skill levels and movement patterns specified by input style exemplars. Robo-Saber demonstrates promise in synthesizing rich gameplay data for predictive applications and enabling a physics-based whole-body VR playtesting agent.
Figures & tables
Figure 1: Overview of our generative model architecture, extending Categorical Codebook Matching (CCM, [ 49 ] , Section 3 ) with Transformer encoder models Estyle and Egame , as well as a modified loss function based on Jensen-Shannon divergence.
Figure 2: Comparing Robo-Saber’s performance to that of human players. Robo-Saber trajectories are produced by using Nref=5 segments from elite (top 5%) players. Left: Human and Robo-Saber TS score distributions across all held-out maps are shown as raincloud plots. Right: Robo-Saber’s performance relative to human players is quantified as score percentiles. To compute the percentiles, for each difficulty level, we select the 100 most-played maps (400 total) and compare Robo-Saber’s score against human scores in each map. The number of human scores available in each map is visualized by opaqueness (max=388, min=3) The performance statistics are summarized as mean ± standard error. The reference-aware model with 5 reference segments from top players ( Nref=5 ) outperforms humans on average.
Figure 3: Sampling-based candidate trajectory selection improves performance compared to using deterministic Argmax selection for GS-VAE during inference. The percentiles are evaluated on the same held-out maps as Fig. 2 with Nref=5 and x:Nref sampled from elite (top 5%) players. The TS score percentiles are summarized as mean ± standard error for each configuration.
Figure 4: Quantifying Robo-Saber’s ability to produce trajectories consistent with the style reference. The oracle player classifier’s top- k accuracy measures how well the generated 3p trajectories are recognized. Adding style reference segments clearly improves the recognizability.
Figure 5: Quantifying Robo-Saber variants’ calibration to the skill levels of human players, measured by the Pearson correlation ( r ) between Robo-Saber and human players’ performance on held-out maps. Left: Densities of points comparing robot (predicted) scores against ground truth human scores. The reference-aware Robo-Saber with Nref=5 achieves a strong correlation of r=0.789 . Right: Distributions of score differences between humans and Robo-Saber variants.
Figure 6: Left: 2D density plot comparing predicted and ground truth TS scores evaluated on N . We compare our FM configuration (violet) with the direct player simulation results (green). A significant reduction in MSE and improvement in Pearson’s r is detected. Right: A raincloud plot comparing the residuals. We report the MSE ± standard error in the legend.
Figure 7: Evaluating the performances of kinematic (left) and physics-based (right) versions of Robo-Saber ( Nref=5 for both). The distribution of Robo-Saber’s TS score percentiles per map is shown across difficulty levels. Physics-based trajectories are produced by tracking the generated 3p trajectories. As expected, absolute performance degrades due to physical/embodied constraints and tracking errors. The kinematic agent achieves respectable percentiles across all difficulty levels. While the physics-based agent achieves percentiles above 40% on average on Normal maps, the performance is poor on Expert and above.
Figure 8: Examining the effect of including physics-based tracking results in factorization machine (FM) training. Each point corresponds to a converged FM model’s report of the validation metric in question. K=kinematic, P=physics-based, KP=both kinematic and physics-based, annotating the synthetic dataset included in FM training. “Mean” is the baseline of using the mean score in R as a constant predictor. Standard error of the mean is visualized with bars at each point. Left: As the FM embedding size grows, MSE on N for KP converges to lower values than those of others, while possibly underfitting at smaller embedding sizes. P nearly competes with K, suggesting the presence of some meaningful signal for PSP. The inset figure zooms into the curves at the bottom. Right: Pearson’s r for N converges to higher values for KP at larger embedding sizes but shows diminishing returns.
Figure 9: The physics-based version of Robo-Saber plays colored notes. The red and blue notes are correctly cut in sequence, with matching saber colors and directions. The dotted notes can be hit from any direction.
Figure 10: Robo-Saber avoids a long sequence of bomb notes by moving its hands up and down and orienting the sabers away from them.
Figure 11: Robo-Saber avoids obstacles (red boxes) by ducking to lower its head and swaying away from them. Robo-Saber is first in performing articulated whole-body gameplay movements for VR.
Figure 12: Robo-Saber samples viable 3p trajectory samples (shown as semi-transparent headset and sabers) for the same game state input, from which the most optimal one is selected. Many possible trajectories for cutting the colored notes appear, and after evaluation, the one with the largest pre- and post-swing angle is selected.
Figure 13: Top: Conditioning Robo-Saber on different x:Nref induces variations in behavior. Some correct and incorrect saber movements emerge, reflecting skill level variations present in conditioning signals. Bottom: Varying the random seed instead while conditioned on the same reference, the movement variation is less noticeable.
Figure 14: Two image sequences, each comparing input reference segments (first 5 rows) and the resulting 3p generation (last row). Top: Conditioning on an expert’s skilled gameplay results in finessed and confident movements, featuring anticipation and fast swing speed. Bottom: Conditioning on a novice’s reference segments yields hesitant and cautious positioning. The two are evaluated on the same map.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Symbol
Description
Value
T
Number of frames per 3p motion chunk
16
-
Interval for interpolating the T frames
4
n
Length of latent sequence for each domain (colored notes, bomb notes, obstacles)
20
s
The default lookahead, in seconds, used in training
2.0
h
Number of frames used as 3p history input
2
λMatch
Jensen-Shannon divergence-based matching loss weight
1e-4
Appendix
Table 1: Hyperparameters and their values used for our experiments.
School of Computer Engineering and Science, Shanghai University, Shanghai, 200444, China · School of Future Technology, Shanghai University, Shanghai, 200444, China