Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective
Organizations: Department of Biomedical Engineering and Sports Technology The Hong Kong Polytechnic University
Abstract
Medical time series (MedTS) underpin many clinical classification tasks, yet existing methods usually represent them only as numerical sequences and underuse the morphology that is explicit in waveform inspection. To bridge this gap, we introduce Vision-Informed Retrieval (ViRe), which uses a frozen VLM-derived waveform representation as a morphology-aware Query to guide retrieval from raw numerical MedTS features. Specifically, a Vision Query is extracted using pre-trained vision-language models (VLMs) to obtain morphology-aware priors from waveform plots. A tailored attention-based cross-modal retrieval mechanism then uses the Vision Query to select morphology-relevant temporal and channel evidence from the numerical representation. ViRe demonstrates strong effectiveness against ten established baselines, yielding an overall 6.42% relative improvement over the previous state of the art across six public benchmarks. Code, training scripts, and reproducibility materials are publicly available in the GitHub Repository: https://github.com/Levi-Ackman/ViRe.
Figures & tables
| ADFTD | APAVA | PTB | MIMIC | ||||||
| Query Type | Accuracy | F1-Score | Accuracy | F1-Score | Accuracy | F1-Score | Accuracy | F1-Score | Avg. Gain |
| w/o | 54.79 1.81 | 51.73 2.22 | 83.86 2.87 | 82.45 3.69 | 81.45 3.44 | 74.58 6.89 | 84.92 1.11 | 84.81 1.12 | – |
| Zero | 54.72 2.82 | 51.23 2.72 | 84.34 1.72 | 83.73 1.96 | 84.62 2.55 | 80.21 3.96 | 86.13 0.14 | 86.06 0.13 | 1.92% |
| Gaussian | 54.46 2.21 | 51.54 1.93 | 82.89 2.66 | 81.65 3.12 | 82.83 2.15 | 77.34 3.64 | 86.50 0.11 | 86.44 0.11 | 0.76% |
| Vision | 57.83 1.82 | 54.02 2.22 | 91.43 1.03 | 91.19 1.03 | 88.26 1.12 | 85.81 1.52 | 88.61 0.12 | 88.54 0.12 | 7.72% |
| ADFTD | APAVA | PTB | MIMIC | ||||||
| Fusion Strategy | Accuracy | F1-Score | Accuracy | F1-Score | Accuracy | F1-Score | Accuracy | F1-Score | Avg. Gain |
| w/o | 54.79 1.81 | 51.73 2.22 | 83.86 2.87 | 82.45 3.69 | 81.45 3.44 | 74.58 6.89 | 84.92 1.11 | 84.81 1.12 | – |
| Add | 56.67 1.46 | 53.53 2.12 | 85.95 1.77 | 85.45 1.71 | 84.81 2.17 | 81.26 3.47 | 87.57 0.06 | 87.49 0.07 | 4.05% |
| Concat | 56.86 1.54 | 53.55 1.21 | 84.23 2.55 | 83.02 3.11 | 83.60 2.25 | 79.13 3.47 | 87.88 0.18 | 87.81 0.18 | 3.02% |
| Retrieval | 57.83 1.82 | 54.02 2.22 | 91.43 1.03 | 91.19 1.03 | 88.26 1.12 | 85.81 1.52 | 88.61 0.12 | 88.54 0.12 | 7.72% |
| ADFTD | APAVA | PTB | MIMIC | ||||||
| Vision Backbone | Accuracy | F1-Score | Accuracy | F1-Score | Accuracy | F1-Score | Accuracy | F1-Score | Avg. Gain |
| w/o | 54.79 1.81 | 51.73 2.22 | 83.86 2.87 | 82.45 3.69 | 81.45 3.44 | 74.58 6.89 | 84.92 1.11 | 84.81 1.12 | – |
| ViT (Random Init) | 37.39 3.11 | 27.34 1.92 | 80.62 0.60 | 77.18 0.54 | 80.65 1.68 | 73.50 2.79 | 83.67 0.14 | 83.58 0.13 | -11.81% |
| ViT (ImageNet) | 52.21 0.73 | 50.55 1.13 | 84.48 2.80 | 84.86 3.67 | 82.06 3.54 | 75.72 5.83 | 86.52 0.21 | 86.42 0.21 | 0.34% |
| CLIP-Vision | 57.83 1.82 | 54.02 2.22 | 91.43 1.03 | 91.19 1.03 | 88.26 1.12 | 85.81 1.52 | 88.61 0.12 | 88.54 0.12 | 7.72% |
| ADFTD | PTB | MIMIC | ||||||||||
| Metrics/Features | Vision | Zero | Gaussian | Raw | Vision | Zero | Gaussian | Raw | Vision | Zero | Gaussian | Raw |
| DBI | 4.034 | 7.962 | 18.098 | 9.607 | 1.156 | 1.466 | 1.722 | 1.385 | 0.893 | 1.216 | 32.475 | 1.562 |
| NMI | 0.101 | 0.054 | 0.004 | 0.042 | 0.425 | 0.345 | 0.117 | 0.194 | 0.398 | 0.327 | 0.002 | 0.276 |
| Homogeneity | 0.112 | 0.055 | 0.004 | 0.040 | 0.441 | 0.371 | 0.122 | 0.202 | 0.403 | 0.332 | 0.002 | 0.274 |
| Completeness | 0.102 | 0.053 | 0.003 | 0.039 | 0.411 | 0.332 | 0.113 | 0.186 | 0.401 | 0.329 | 0.003 | 0.281 |
| MAR | CAR | |||
| Dataset | ViRe | Gaussian | ViRe | Gaussian |
| PTB | 1.81 | 1.46 | 1.65 | 1.33 |
| PTB-XL | 1.67 | 1.38 | 1.51 | 1.27 |
| MIMIC | 1.72 | 1.41 | 1.63 | 1.35 |
| A. Diagnosis F1 | ||||
| Diagnosis | ViRe | Medformer | Gain | Patient-bootstrap 95% CI |
| MI | 72.21 | 71.35 | +0.85 | [-1.42,+3.08] |
| STTC | 74.91 | 73.44 | +1.47 | [-0.99,+3.97] |
| HYP | 59.49 | 44.17 | +15.32 | [+10.39,+20.25] |
| AMI | 72.93 | 71.01 | +1.92 | [-1.23,+5.14] |
| ISCA | 37.07 | 30.15 | +6.92 | [-2.03,+15.94] |
| Attribute | Metric | CLIP | Random | Gain |
| Amplitude range | 0.903 | 0.501 | +0.402 | |
| PR interval | 0.633 | 0.248 | +0.384 | |
| Cross-lead synchrony | 0.684 | 0.320 | +0.364 | |
| Prolonged QT | AUROC | 0.904 | 0.656 | +0.248 |
| Zero-shot text alignment | 12/12 directions |
| Key results | AF AUROC +0.423 ; QTc +0.284 ; amplitude +0.239 ; synchrony +0.232 . |
| Same-patient natural changes | 5/5 measurements |
| Key results | Direction accuracy: HR +0.128 ; PR +0.154 ; QRS +0.127 ; QT +0.238 ; QTc +0.144 . |
| Raw-space retrieval | 21/21 positive CIs |
| Key results | Bradycardia +0.337 ; HR +0.331 ; amplitude +0.262 ; QT/QRS +0.256/+0.245 . |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Operation | Transformation | Default |
| Temporal flipping | Reverse the sequence along the time axis. | |
| Channel shuffling | Randomly permute the channel order. | |
| Temporal masking | Mask timestamps shared across all channels. | |
| Frequency masking | Suppress randomly selected frequency bands and transform the signal back to the time domain. | |
| Jittering | Add random noise sampled from and scale its magnitude. | |
| Dropout | Randomly set a fraction of signal values to zero. |
| Item | Setting |
| Batch size | |
| Learning rate | |
| Model dimension | |
| Temporal Encoder depth | for all datasets |
| Channel Encoder depth | by default; for TDBrain |
| Temporal granularity | Dataset order: APAVA, TDBrain, ADFTD, PTB, PTB-XL, MIMIC; |
| Metric | ViRe | Medformer | FEDformer |
| F1-Score | 91.19 1.01 | 76.31 0.71 | 73.51 3.30 |
| Training Time (ms) | 211.9 | 313.1 | 477.7 |
| Training Memory (MB) | 698.2 | 838.4 | 1103.9 |
| Inference Time (ms) | 267.7 | 152.6 | 282.6 |
| Inference Memory (MB) | 740.8 | 408.5 | 491.6 |
| Evaluation | Gain | Wins | Statistical evidence |
| PTB-XL, 13 diagnoses (64.93 vs. 61.57) | +3.36 | n/a | Cluster 95% CI [+1.46,+5.10] ; paired . |
| PTB-XL, five superclasses (72.26 vs. 68.52) | +3.74 | n/a | Cluster 95% CI [+2.36,+5.14] ; paired . |
| PTB-XL, 13 diagnoses; repeated splits | +2.46 | 7/7 | Wilcoxon ; corrected paired -test ; 95% CI [+1.14,+3.77] . |
| PTB-XL, five superclasses; repeated splits | +2.42 | 7/7 | Wilcoxon ; corrected paired -test ; 95% CI [+0.99,+3.84] . |
| APAVA; repeated splits | +6.48 | 8/10 | Exact Wilcoxon/Holm . |
| PTB; repeated splits | +3.98 | 9/10 | Exact Wilcoxon/Holm . |
| Dataset | Metric | ViRe | Hi-Patch | t-PatchGNN | GRU-D |
| Human Activity | MSE | 2.45 0.04 | 2.57 0.02 | 2.66 0.03 | 3.94 0.29 |
| MAE | 3.04 0.02 | 3.11 0.03 | 3.15 0.02 | 4.37 0.21 | |
| PhysioNet | MSE | 4.56 0.04 | 4.86 0.03 | 4.98 0.08 | 5.76 0.34 |
| MAE | 3.44 0.03 | 3.62 0.07 | 3.72 0.03 | 4.53 0.15 | |
| MIMIC-III | MSE | 1.56 0.09 | 1.75 0.26 | 1.69 0.03 | 2.35 0.06 |
| MAE | 6.77 0.11 | 7.24 0.18 | 7.22 0.09 | 8.34 0.22 |
| Dataset | Metric | ViRe | MTM | STraTS | t-PatchGNN | GRU-D |
| P12 | AUROC | 88.5 1.2 | 88.0 1.0 | 86.4 1.1 | 84.5 0.9 | 81.9 2.1 |
| AUPRC | 60.3 2.4 | 58.6 4.1 | 53.9 3.1 | 50.8 2.6 | 46.1 4.7 | |
| P19 | AUROC | 91.3 2.0 | 90.3 2.0 | 89.7 1.8 | 87.0 1.4 | 83.9 1.7 |
| AUPRC | 62.3 4.3 | 58.3 5.3 | 57.9 3.3 | 51.5 5.2 | 46.9 2.1 | |
| PAM | Accuracy | 98.3 0.6 | 97.5 0.2 | 96.4 0.8 | 93.9 1.2 | 83.3 1.6 |
| F1 | 98.4 0.6 | 97.6 0.2 | 95.3 0.7 | 94.8 1.2 | 84.8 1.2 |