When Minute-Resolution Monitoring Meets Session-Level Injury Labels: Landmark-Based Discrimination in Elite Women's Football
Abstract
Minute-resolution athlete monitoring is increasingly common, while injury annotation may exist only at the athlete-session level and omit within-session onset time. Replicating a positive session label across every recorded minute would therefore create unsupported minute-level supervision. We address this label-resolution mismatch using fixed elapsed-time landmarks at 10, 20, 30, 40, 50, and 60 min, constructing one representation per athlete-session from information available up to each landmark while keeping the target as a same-day injury-associated session indicator. Using 2020 SoccerMon data from elite women's football, the modelling cohort contains 2,259 Team A athlete-sessions from 27 athletes, including all 22 positive sessions from five athletes. Evaluation is athlete-disjoint. We compare contextual, cumulative, and dynamic representations; Logistic Regression, Random Forest, XGBoost, and TabPFN; and training-only NONE, SMOTE, and CTGAN conditions. Robustness is assessed using athlete-cluster bootstrap, a fixed common cohort, alternative fold allocations, leave-one-positive-athlete-out analysis, and equal-athlete weighting. Discrimination is landmark-dependent and non-monotonic. TabPFN improves later-landmark discrimination relative to Logistic Regression but does not consistently outperform Random Forest. Synthetic augmentation provides condition-specific rather than universal benefit. The contribution is a unit-aligned framework for using minute-resolution predictors with session-level supervision, not minute-specific injury prediction.