Semantic-driven time-series generation offers a promising way to improve downstream learning in few-shot forecasting, but directly generating numerical sequences from language often fails to preserve the intended temporal structure. We propose VisualBridge, which uses time-series plots as a visual intermediate to bridge high-level temporal semantics and numerical sequences. An MLLM first converts plotted series into structured semantic representations, enabling explicit control over temporal properties such as trend, seasonality, and volatility. We then learn a semantic editing policy with downstream forecasting rewards, allowing the generation process to favor temporal patterns that are beneficial for the target task. The resulting sequences are further modeled by a temporal VAE to produce consistent multivariate augmentations. Experiments on standard public forecasting benchmarks demonstrate that VisualBridge improves few-shot forecasting over conventional augmentation methods, with ablations validating the roles of visual semantic grounding, learned semantic control, and VAE-based generation.
Figures & tables
Figure 1: Direct numerical generation vs. visual generation for time series. The two right-hand cases use GPT-5.6-Luna for direct numerical generation and GPT-image-2 for plot generation, respectively, showing that the visual bridge better preserves the intended temporal semantics.
Figure 2: Visually-grounded controllable generation pipeline of VisualBridge. Time-series plots connect structured temporal descriptions to numerical generation. An RL policy learns semantic edits using rewards driven by forecasting gains over no-edit generation, measured on held-out probes within the few-shot training split. Edited and no-edit branches use matched initialization and short-run budgets to train forecasters on real and generated data. Generated plots are converted to numerical sequences, and accepted sequences anchor a temporal VAE for multivariate augmentation.
Figure 3: Qualitative comparison between direct generation and VisualBridge. Examples use identical, unmodified semantic records across six benchmarks. VisualBridge uses a plot intermediate to preserve recurring patterns and local shape in the recovered numerical sequence.
Dataset
Per-F1 ↑
Cyc-MAE ↓
Peak-NLE ↓
DF-Err. ↓
TPR-Err. ↓
PE-Err. ↓
D
VB
D
VB
D
VB
D
VB
D
VB
D
VB
ETTh1
64.5%
84.3%
3.6190
2.5794
0.2468
0.2303
0.0381
0.0200
0.1731
0.0541
0.0981
0.0614
ETTh2
62.3%
72.7%
4.0000
2.9032
0.2671
0.2239
0.0399
0.0240
0.2171
0.0768
0.1538
0.1331
ETTm1
55.5%
62.0%
1.5939
1.4789
0.2287
0.2283
0.0030
0.0018
0.1565
0.0629
0.0909
0.0622
ETTm2
69.7%
69.0%
1.5699
1.4956
0.2233
0.2088
0.0060
0.0040
0.2003
0.0638
0.2263
0.2002
Exchange
26.3%
27.0%
1.5161
1.5484
0.0969
0.0939
0.0054
0.0023
0.1411
0.0569
0.1410
0.0975
Table 1: Semantic agreement and temporal-structure preservation. Direct (D) and VisualBridge (VB) generate from the same semantic records. Per-F1 evaluates the requested seasonality; the remaining metrics compare generated and source structure. Better results within each pair are bold.
Dataset
Original
Gaussian
Convolve
TimeGAN
ADA
VisualBridge
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
ETTh1
0.459
0.472
0.463
0.478
0.465
0.477
0.471
0.480
0.460
0.472
0.442 ± 0.001
0.459 ± 0.001
ETTh2
0.353
0.307
0.356
0.308
0.358
0.311
0.357
0.314
0.359
0.318
0.334 ± 0.002
0.292 ± 0.001
ETTm1
0.416
0.428
0.414
0.427
0.404
0.433
0.407
0.421
0.406
0.426
0.403 ± 0.001
0.415 ± 0.002
ETTm2
0.268
0.190
0.269
0.190
0.272
0.193
0.273
0.192
0.270
0.192
0.263 ± 0.001
0.186 ± 0.001
Weather
0.212
0.176
0.220
0.184
0.232
0.192
0.219
0.180
0.224
0.185
0.209 ± 0.001
0.173 ± 0.001
Table 2: Main results on few-shot forecasting. Results for stochastic methods are averaged over three generated datasets under a fixed forecasting seed. VisualBridge reports mean ± standard deviation and consistently outperforms existing data augmentation baselines across all eight public benchmarks.
Dataset
Gaussian
Convolve
TimeGAN
ADA
VisualBridge
FMAE
FMSE
FMAE
FMSE
FMAE
FMSE
FMAE
FMSE
FMAE
FMSE
ETTh1
-13.0%
-21.8%
-19.6%
-18.1%
-39.1%
-29.0%
-3.3%
0.0%
55.4%
47.2%
ETTh2
-25.6%
-5.5%
-42.7%
-21.9%
-34.2%
-38.4%
-51.3%
-60.3%
162.4%
82.3%
ETTm1
3.4%
0.9%
20.1%
-4.3%
15.1%
6.0%
16.8%
1.7%
21.8%
11.1%
ETTm2
-10.5%
0.0%
-42.1%
-17.9%
-52.6%
-11.9%
-21.0%
-11.9%
52.6%
23.9%
Weather
-72.6%
-94.4%
-181.6%
-188.9%
-63.6%
-47.2%
-109.0%
-106.3%
27.2%
35.4%
Table 3: Performance gap recovery measured by F . Percentage of the few-shot-to-full-data error gap recovered by augmentation. Values above 100% surpass the full-data reference; negative values indicate degradation relative to the unaugmented few-shot baseline.
Figure 4: Core component ablation by dataset. Bars show mean forecasting MSE. Random Semantic and w/o VAE both omit VAE sampling and differ in whether the semantic action sampler is learned. Each panel uses a dataset-specific, nonzero MSE-axis lower bound; error bars are omitted.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
ETTh1
ETTh2
ETTm1
ETTm2
Exchange
Weather
Electricity
Traffic
D
VB
D
VB
D
VB
D
VB
D
VB
D
VB
D
VB
D
VB
Tr-Acc. ↑
84.9%
88.9%
77.0%
75.4%
93.5%
91.7%
87.5%
86.5%
76.8%
71.4%
84.7%
84.7%
60.6%
60.5%
79.5%
77.6%
Tr-F1 ↑
46.9%
58.4%
51.8%
48.7%
45.6%
47.2%
48.1%
41.2%
56.3%
53.9%
69.7%
67.0%
35.2%
46.1%
32.1%
28.1%
Cyc-Acc. ↑
11.9%
27.0%
16.1%
29.8%
50.2%
52.1%
53.1%
57.0%
61.3%
61.3%
67.8%
69.7%
35.5%
36.1%
28.4%
20.9%
Inc.- W1↓
0.1104
0.0583
0.1483
0.0906
0.0454
0.0439
0.1101
0.1118
0.0348
0.0302
0.0750
0.1197
0.1318
0.0859
0.1163
0.0714
BP- ℓ1↓
0.3464
0.2535
0.3696
0.2831
0.1049
0.1038
0.1626
0.1789
0.1116
0.0783
0.0899
0.0892
0.5296
0.4745
0.4424
0.3445
Appendix
Table 4: Complementary semantic and structural diagnostics. Direct (D) and VisualBridge (VB) use the same paired windows as Table 1 .
Metric
ETTh1
Exchange
Direct
VisualBridge
Direct
VisualBridge
Tr-Acc. ↑
74.6%
87.3%
83.9%
85.7%
Tr-F1 ↑
32.6%
46.3%
84.5%
64.3%
Per-F1 ↑
71.6%
84.3%
12.5%
21.1%
Cyc-Acc. ↑
15.9%
33.3%
41.9%
51.6%
Cyc-MAE ↓
3.4683
2.6667
1.6774
1.5806
Appendix
Table 5: Same-model comparison between direct numerical generation and VisualBridge. ETTh1 and Exchange contain 126 and 56 paired windows, respectively. Both pathways use gemini-3.1-flash-image-preview with identical semantic records, so the comparison isolates the output pathway. The cycle metrics use 126 ETTh1 and 31 Exchange windows.
Dataset
Trend
Seasonality
Volatility
ETTh1
85.9
79.2
100.0
ETTh2
89.4
81.5
77.3
ETTm1
94.2
74.5
99.1
ETTm2
78.4
85.2
92.4
Weather
82.8
73.6
100.0
Electricity
83.4
80.0
76.9
Appendix
Table 6: Realization of semantic actions in matched generations. Cells report the percentage of paired generations in which the requested attribute changes in the intended direction.
Dataset
Matched identity
Original-only reference
ETTh1
0.459 ± 0.001
0.466 ± 0.001
ETTh2
0.292 ± 0.001
0.299 ± 0.001
ETTm1
0.415 ± 0.002
0.421 ± 0.002
ETTm2
0.186 ± 0.001
0.188 ± 0.001
Weather
0.173 ± 0.001
0.174 ± 0.001
Electricity
0.171 ± 0.000
0.172 ± 0.000
Appendix
Table 7: Effect of the task-utility reference. Test MSE of final forecasters trained with augmentation from each policy.
Backbone
Dataset
Original
VisualBridge
MAE
MSE
MAE
MSE
PatchTST
ETTh1
0.484
0.512
0.465
0.496
PatchTST
ETTm1
0.405
0.416
0.389
0.397
PatchTST
Weather
0.213
0.178
0.206
0.174
PatchTST
Exchange
0.211
0.091
0.207
0.088
DLinear
ETTh1
0.460
0.469
0.445
0.451
Appendix
Table 8: Augmentation transfer across forecasting backbones. Original uses real data only; VisualBridge adds a fixed augmented set.
Dataset
Features
Total Windows
Standard Partition
Few-shot Partition
ETTh1 / ETTh2
7
14019
8449 / 2785 / 2785
1689 / 2785 / 2785
ETTm1 / ETTm2
7
57219
34369 / 11425 / 11425
6873 / 11425 / 11425
Traffic
862
17163
12089 / 1661 / 3413
2417 / 1661 / 3413
Electricity
321
25923
18221 / 2537 / 5165
3644 / 2537 / 5165
Weather
21
52315
36696 / 5175 / 10444
7339 / 5175 / 10444
Exchange
8
7207
5120 / 665 / 1422
1024 / 665 / 1422
Appendix
Table 9: Dataset statistics and partitions. Counts refer to windows with 96 input and 96 prediction steps; partitions are train/validation/test. Few-shot training uses the earliest 20% of training windows.
Component
Setting
Value
Forecasting and visual generation
Forecaster
Backbone and objective
iTransformer; L1 loss
Forecaster
Hidden dimensions
dmodel=128 , dff=128
Forecaster
Encoder architecture
2 layers; 8 attention heads
Forecaster
Optimization
Adam; learning rate 10−4
Semantic QA
Model
gpt-5.6-luna
Appendix
Table 10: Model and forecasting settings. All final forecasters use seed 2023 ; repeated results vary the augmented data.
Component
Setting
Value
Action sampling and policy updates
Action sampling
Group size and exploration
4 actions; mixture weight 0.20
Candidate blending
Generated-sequence weight ρ
0.10
Policy update
Optimizer and steps per group
Adam; 12 steps
Policy update
Learning rate
0.03
Policy update
Clipping and regularization
Clip 0.20 ; KL 0.08 ; entropy 0.02
Appendix
Table 11: Semantic action sampler and reward settings.