LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
Organizations: University of North Carolina at Chapel Hill · Foci Labs
Abstract
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the described time series from multiple distractors. Matching is a far lighter demand on the verifier than writing questions or judging a caption, so an off-the-shelf LLM can supply the reward. Across two captioning benchmarks, and on forecasting and reconstruction where the predictor sees only the caption, LineupRL outperforms SFT and RL baselines on every metric. The 3B vision language model (VLM) trained by LineupRL also outperforms, at 1/24 of the parameters, the 72B VLM whose captions the SFT baseline is distilled from. Our case study shows that LineupRL resists reward hacking, and that the captioner it trains both traces the trend and names the values at key points.
Figures & tables
| BEDTime | CaTS-Bench | ||||||
| Type | Captioner | Entailment | c s | s c | Entailment | c s | s c |
| VLM | Qwen2.5-VL-3B | 0.296 | 0.550 | 0.495 | 0.352 | 0.751 | 0.721 |
| Qwen2.5-VL-72B | 0.320 | 0.605 | 0.484 | 0.440 | 0.803 | 0.643 | |
| Qwen3-VL-8B | 0.320 | 0.669 | 0.498 | 0.407 | 0.817 | 0.626 | |
| InternVL3-14B | 0.357 | 0.562 | 0.512 | 0.506 | 0.766 | 0.701 | |
| LLM | Phi-3.5-mini | 0.180 | 0.909 | 0.807 | 0.318 | 0.980 | 0.890 |
| Forecasting | Reconstruction | |||||||
| Captioner | ETTh2 | ETTm2 | Saug. | Elec. | ETTh2 | ETTm2 | Saug. | Elec. |
| Naive baseline | 0.3047 | 0.0596 | 1.2955 | 0.9078 | 0.9501 | 0.8335 | 1.1748 | 0.8634 |
| Empty caption | 0.8446 [-2.5pt] 0.0101 | 0.8001 [-2.5pt] 0.0042 | 1.0601 [-2.5pt] 0.0016 | 0.9129 [-2.5pt] 0.0041 | 0.9284 [-2.5pt] 0.0306 | 0.8274 [-2.5pt] 0.0029 | 1.1764 [-2.5pt] 0.0008 | 0.8667 [-2.5pt] 0.0021 |
| Shuffled caption | 0.8437 [-2.5pt] 0.0087 | 0.8040 [-2.5pt] 0.0046 | 1.0619 [-2.5pt] 0.0042 | 0.9124 [-2.5pt] 0.0043 | 0.9347 [-2.5pt] 0.0121 | 0.8270 [-2.5pt] 0.0015 | 1.1774 [-2.5pt] 0.0018 | 0.8663 [-2.5pt] 0.0010 |
| Qwen2.5-VL-3B | 0.3062 [-2.5pt] 0.0086 | 0.1012 [-2.5pt] 0.0047 | 0.9419 [-2.5pt] 0.0119 | 0.5751 [-2.5pt] 0.0351 | 0.2445 [-2.5pt] 0.0052 | 0.0633 [-2.5pt] 0.0029 | 0.6852 [-2.5pt] 0.0173 | 0.4352 [-2.5pt] 0.0156 |
| Qwen2.5-VL-72B | 0.2149 [-2.5pt] 0.0094 | 0.0451 [-2.5pt] 0.0009 | 0.8521 [-2.5pt] 0.0061 | 0.3089 [-2.5pt] 0.0094 | 0.1134 [-2.5pt] 0.0065 | 0.0288 [-2.5pt] 0.0022 | 0.4444 [-2.5pt] 0.0113 | 0.1993 [-2.5pt] 0.0077 |
| Nearest | Random | |
| Distractors | ( ours ) | |
| Reward acc. over chance | 0.872 | 0.982 |
| Entailment | 0.506 | 0.354 |
| Ident. c s | 0.797 | 0.704 |
| Ident. s c | 0.753 | 0.744 |
| Caption length (chars) | 667 | 738 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| LineupRL | RL, answerability | RL, LLM-as-judge | SFT | |
| Training signal | identification among candidates | accuracy on the chart’s verified questions | grade – of the caption | cross-entropy on 72B captions |
| Frozen model in the loop | Qwen2.5-14B-Instruct verifier | Qwen2.5-14B-Instruct answerer | Qwen2.5-14B-Instruct judge | Qwen2.5-VL-72B-Instruct teacher |
| Training charts | ( questions) | pairs | ||
| Passes over the data | ||||
| Learning rate | , cosine | |||
| Checkpoint every | steps | steps | steps | steps |
| SFT | Answ. | Judge | 72B | 3B init | |
| Forecasting | |||||
| ETTh2 | 0.001 | 0.001 | 0.002 | 0.002 | 0.001 |
| ETTm2 | 0.016 | 0.001 | 0.035 | 0.001 † | 0.001 |
| Saug. | 0.007 | 0.001 | 0.030 | 0.005 | 0.001 |
| Elec. | 0.001 | 0.001 | 0.001 | 0.001 | 0.001 |
| Reconstruction | |||||
| Forecasting (predict the next window) | ||||
| Captioner | ETTh2 | ETTm2 | Saug. | Elec. |
| Naive baseline | 0.3047 | 0.0596 | 1.2955 | 0.9078 |
| Empty caption | 0.8446 0.0101 | 0.8001 0.0042 | 1.0601 0.0016 | 0.9129 0.0041 |
| Shuffled caption | 0.8437 0.0087 | 0.8040 0.0046 | 1.0619 0.0042 | 0.9124 0.0043 |
| Qwen2.5-VL-3B | 0.3062 0.0086 | 0.1012 0.0047 | 0.9419 0.0119 | 0.5751 0.0351 |
| Qwen2.5-VL-72B | 0.2149 0.0094 | 0.0451 0.0009 | 0.8521 0.0061 | 0.3089 0.0094 |
| Qwen2.5-14B | GLM-4-9B | Llama-3.1-8B | Mistral-Nemo-12B | Phi-4-14B | ||||||
| Captioner | c s | s c | c s | s c | c s | s c | c s | s c | c s | s c |
| Qwen2.5-VL-3B | 0.664 | 0.566 | 0.559 | 0.443 | 0.472 | 0.328 | 0.518 | 0.349 | 0.541 | 0.548 |
| Qwen2.5-VL-72B | 0.742 | 0.625 | 0.610 | 0.455 | 0.582 | 0.338 | 0.574 | 0.334 | 0.600 | 0.513 |
| Qwen3-VL-8B | 0.752 | 0.604 | 0.639 | 0.446 | 0.648 | 0.372 | 0.620 | 0.351 | 0.700 | 0.549 |
| InternVL3-14B | 0.709 | 0.600 | 0.585 | 0.467 | 0.542 | 0.360 | 0.541 | 0.324 | 0.539 | 0.557 |
| Phi-3.5-mini | 0.936 | 0.903 | 0.893 | 0.828 | 0.884 | 0.668 | 0.837 | 0.701 | 0.925 | 0.787 |
| Qwen2.5-14B | GLM-4-9B | Llama-3.1-8B | Mistral-Nemo-12B | Phi-4-14B | ||||||
| Captioner | c s | s c | c s | s c | c s | s c | c s | s c | c s | s c |
| Qwen2.5-VL-3B | 0.873 | 0.803 | 0.674 | 0.680 | 0.675 | 0.429 | 0.591 | 0.449 | 0.827 | 0.763 |
| Qwen2.5-VL-72B | 0.951 | 0.806 | 0.711 | 0.631 | 0.763 | 0.372 | 0.648 | 0.434 | 0.896 | 0.655 |
| Qwen3-VL-8B | 0.933 | 0.775 | 0.716 | 0.587 | 0.802 | 0.379 | 0.650 | 0.413 | 0.919 | 0.665 |
| InternVL3-14B | 0.919 | 0.798 | 0.684 | 0.689 | 0.734 | 0.384 | 0.619 | 0.426 | 0.849 | 0.713 |
| Phi-3.5-mini | 0.989 | 0.988 | 0.979 | 0.953 | 0.978 | 0.819 | 0.867 | 0.849 | 0.981 | 0.826 |
| Domain | Verifier | |||||
| TSFragment val. | Qwen2.5-14B | 0.701 | 0.634 | 0.586 | 0.527 | 0.496 |
| Qwen2.5-7B | 0.641 | 0.553 | 0.491 | 0.443 | 0.401 | |
| Qwen2.5-3B | 0.455 | 0.353 | 0.284 | 0.186 | 0.233 | |
| CaTS-Bench | Qwen2.5-14B | 0.421 | 0.384 | 0.357 | 0.336 | 0.324 |
| Qwen2.5-7B | 0.257 | 0.199 | 0.170 | 0.151 | 0.147 | |
| Qwen2.5-3B | 0.112 | 0.100 | 0.095 | 0.080 | 0.105 |
| BEDTime | CaTS-Bench | |||||||
| Reward model | Ent. | c s | s c | Len. | Ent. | c s | s c | Len. |
| Untuned Qwen2.5-VL-3B | 0.296 | 0.550 | 0.495 | 1084 | 0.352 | 0.751 | 0.721 | 1287 |
| Qwen2.5-14B ( ours ) | 0.400 | 0.709 | 0.638 | 621 | 0.611 | 0.886 | 0.868 | 713 |
| Qwen2.5-7B | 0.287 | 0.710 | 0.644 | 1331 | 0.298 | 0.850 | 0.799 | 1725 |
| Qwen2.5-3B | 0.270 | 0.715 | 0.710 | 898 | 0.381 | 0.830 | 0.831 | 959 |
| Forecasting | Reconstruction | |||||||||
| Reward model | ETTh2 | ETTm2 | Saug. | Elec. | Rel. | ETTh2 | ETTm2 | Saug. | Elec. | Rel. |
| Untuned Qwen2.5-VL-3B | 0.3114 | 0.1037 | 0.9234 | 0.5476 | 1.850 | 0.2469 | 0.0654 | 0.6698 | 0.4124 | 2.283 |
| Qwen2.5-14B ( ours ) | 0.1687 | 0.0523 | 0.8225 | 0.1922 | 1.000 | 0.0974 | 0.0286 | 0.4472 | 0.1317 | 1.000 |
| Qwen2.5-7B | 0.2535 | 0.0758 | 0.8709 | 0.2926 | 1.369 | 0.1441 | 0.0435 | 0.5318 | 0.1847 | 1.391 |
| Qwen2.5-3B | 0.2312 | 0.0783 | 0.8459 | 0.3205 | 1.370 | 0.1693 | 0.0534 | 0.6552 | 0.2385 | 1.713 |
| BEDTime | CaTS-Bench | |||||||
| Candidates | Ent. | c s | s c | Len. | Ent. | c s | s c | Len. |
| Untuned Qwen2.5-VL-3B | 0.296 | 0.550 | 0.495 | 1084 | 0.352 | 0.751 | 0.721 | 1287 |
| ( ours ) | 0.400 | 0.709 | 0.638 | 621 | 0.611 | 0.886 | 0.868 | 713 |
| 0.343 | 0.664 | 0.644 | 482 | 0.470 | 0.851 | 0.869 | 561 | |
| 0.340 | 0.790 | 0.796 | 555 | 0.391 | 0.852 | 0.876 | 620 | |
| Forecasting | Reconstruction | |||||||||
| Candidates | ETTh2 | ETTm2 | Saug. | Elec. | Rel. | ETTh2 | ETTm2 | Saug. | Elec. | Rel. |
| Untuned Qwen2.5-VL-3B | 0.3114 | 0.1037 | 0.9234 | 0.5476 | 1.850 | 0.2469 | 0.0654 | 0.6698 | 0.4124 | 2.283 |
| ( ours ) | 0.1687 | 0.0523 | 0.8225 | 0.1922 | 1.000 | 0.0974 | 0.0286 | 0.4472 | 0.1317 | 1.000 |
| 0.2046 | 0.0442 | 0.8067 | 0.1925 | 1.002 | 0.0975 | 0.0288 | 0.4179 | 0.1234 | 0.969 | |
| 0.2206 | 0.0529 | 0.8421 | 0.2120 | 1.106 | 0.1006 | 0.0271 | 0.4287 | 0.1382 | 0.996 | |
| BEDTime | CaTS-Bench | |||||||
| Distractors | Ent. | c s | s c | Len. | Ent. | c s | s c | Len. |
| Untuned Qwen2.5-VL-3B | 0.296 | 0.550 | 0.495 | 1084 | 0.352 | 0.751 | 0.721 | 1287 |
| Nearest neighbor ( ours ) | 0.400 | 0.709 | 0.638 | 621 | 0.611 | 0.886 | 0.868 | 713 |
| Random | 0.324 | 0.598 | 0.636 | 723 | 0.384 | 0.810 | 0.853 | 753 |
| Forecasting | Reconstruction | |||||||||
| Distractors | ETTh2 | ETTm2 | Saug. | Elec. | Rel. | ETTh2 | ETTm2 | Saug. | Elec. | Rel. |
| Untuned Qwen2.5-VL-3B | 0.3114 | 0.1037 | 0.9234 | 0.5476 | 1.850 | 0.2469 | 0.0654 | 0.6698 | 0.4124 | 2.283 |
| Nearest neighbor ( ours ) | 0.1687 | 0.0523 | 0.8225 | 0.1922 | 1.000 | 0.0974 | 0.0286 | 0.4472 | 0.1317 | 1.000 |
| Random | 0.2314 | 0.0643 | 0.8138 | 0.2484 | 1.212 | 0.1397 | 0.0268 | 0.4513 | 0.1652 | 1.142 |