GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals
Abstract
Residual vector quantization (RVQ) turns physiological waveforms into compact token sequences, but conventional masked modeling treats every incorrect token as equally costly. We propose GeoRVQ, a coarse-to-fine masked token model whose objective reflects the local response of a frozen waveform decoder. Decoder-induced costs define geometry-aware soft targets and expected distortion, while quantizer-causal prediction follows residual dependencies from coarse to fine levels. In a descriptive aggregate over MIMIC-IV Waveform, VitalDB, and CODE-15%, GeoRVQ increases exact token accuracy from to , reduces decoded distance from to , and increases R-peak F1 from to under matched model and training conditions. Across 45 held-out code substitutions, decoder-induced cost has a Spearman correlation of with realized decoded cost, compared with for Euclidean codeword distance. These results indicate that decoder-aware objectives can improve waveform and event preservation without requiring a large increase in exact token accuracy.