Financial candlestick forecasting is fundamental to quantitative investment, yet it remains exceptionally challenging due to extremely low signal-to-noise ratios and vast heterogeneity across markets and instruments. Existing approaches have largely attempted to introduce deep learning to capture hidden temporal features, but most adopt an auto-regressive formulation, which leads to error accumulation during inference. Meanwhile, general-purpose time-series foundation models are not tailored to the unique structure of k-line data and yield unsatisfactory performance on downstream candlestick forecasting tasks. To tackle these problems, we introduce KiT, a K-line Diffusion Transformer foundation model, and reformulate future prediction as conditional path generation via flow matching: given a historical context window, the model generates an ensemble of plausible future OHLCV trajectories. We pre-train KiT at multiple parameter scales on billions of candlestick bars spanning multiple markets and timescales. Across three markets and seven resolutions, KiT attains a mean return RankIC of 0.057 and a mean volatility RankIC of 0.66, leading at every timescale and outperforming both task-specific financial forecasters and general time-series foundation models. Code will be available at: https://github.com/Luciferbobo/KiT.
Figures & tables
Configuration
1m ↑
5m ↑
15m ↑
30m ↑
1h ↑
2h ↑
1d ↑
Mean ↑
KiT
− .0215
− .0576
− .0526
− .0474
− .0383
− .0651
− .1168
− .0571
Kronos-base
− .0158
− .0426
− .0442
− .0064
− .0128
− .0514
− .1046
− .0342
Kronos-base-FT
− .0192
− .0536
− .0484
− .0408
− .0304
− .0622
− .1101
− .0521
Sundial
− .0148
− .0386
− .0434
− .0082
− .0046
− .0214
− .0837
− .0209
Chronos-Bolt extended
− .0362
− .0088
− .0186
− .0964
− .0068
− .2159
− .0052
− .0461
Chronos-2
− .0168
− .0286
− .0362
− .0184
− .0162
− .0648
− .0556
− .0054
Table 1: Return forecasting by resolution. Evaluated across three markets in the test window. Each entry is the RankIC between the predicted mean terminal return and the realized close-to-close log-return sum.
Configuration
1m ↑
5m ↑
15m ↑
30m ↑
1h ↑
2h ↑
1d ↑
Mean ↑
KiT
− .7599
− .6947
− .6900
− .5919
− .5958
− .6652
− .6225
− .6600
Kronos-base
− .5900
− .4934
− .3955
− .3049
− .4699
− .5281
− .4498
− .4617
Kronos-base-FT
− .6193
− .5478
− .4466
− .3671
− .4974
− .5306
− .4727
− .4974
Sundial
− .6024
− .5477
− .5164
− .4591
− .5691
− .5534
− .4486
− .5281
Chronos-Bolt extended
− .5766
− .4821
− .4113
− .4289
− .3543
− .3797
− .3310
− .4234
Chronos-2
− .4427
− .3944
− .3671
− .4092
− .4264
− .4452
− .2260
− .3873
Table 2: Volatility prediction by resolution. Evaluated across three markets in the test window. Each entry is the RankIC between the predicted mean realized volatility and the realized sample volatility.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
KiT-S
KiT-M
KiT-L
Parameters
31.0M
101.0M
283.7M
Layers
8
13
21
Width
512
768
1,024
Heads
4
6
8
FFN
1,408
2,048
2,816
Steps
50k
50k
50k
Appendix
Table 3: Model configurations.
KiT-S
KiT-M
KiT-L
Parameters
31.0M
101.0M
283.7M
Layers
8
13
21
Width
512
768
1,024
Heads
4
6
8
FFN
1,408
2,048
2,816
Steps
50k
50k
50k
Appendix
Table 3: Model configurations.
Resolution
History
Horizon
1 minute
1,200
120
5 minutes
960
96
15 minutes
800
64
30 minutes
640
40
1 hour
400
20
2 hours
360
20
Appendix
Table 4: Window size.
Configuration
1m ↑
5m ↑
15m ↑
30m ↑
1h ↑
2h ↑
1d ↑
Mean ↑
KiT
− .0235
− .0552
− .0445
− .0330
− .0241
− .0475
− .1158
− .0491
Kronos-base
− .0072
− .0344
− .0368
− .0244
− .0209
− .0186
− .0048
− .0067
Kronos-base-FT
− .0164
− .0506
− .0392
− .0136
− .0125
− .0348
− .0058
− .0247
Sundial
− .0086
− .0074
− .0362
− .0221
− .0152
− .0284
− .1026
− .0315
Chronos-Bolt extended
− .0364
− .0292
− .0246
− .0724
− .0068
− .2223
− .0276
− .0347
Chronos-2
− .0188
− .0383
− .0372
− .0402
− .0841
− .1789
− .0962
− .0161
Appendix
Table 5: Return IC by resolution. Higher is better; bold marks the best value in each column.
Configuration
1m ↑
5m ↑
15m ↑
30m ↑
1h ↑
2h ↑
1d ↑
Mean ↑
KiT
− .7459
− .6749
− .6888
− .6553
− .6621
− .7586
− .5726
− .6797
Kronos-base
− .5654
− .4369
− .3616
− .3175
− .4491
− .5745
− .4377
− .4490
Kronos-base-FT
− .5836
− .4695
− .3884
− .3759
− .4667
− .5736
− .4470
− .4721
Sundial
− .6055
− .5178
− .5233
− .5041
− .5757
− .5138
− .4081
− .5212
Chronos-Bolt extended
− .5830
− .4926
− .4169
− .4907
− .4261
− .4414
− .1562
− .4296
Chronos-2
− .4156
− .3688
− .3185
− .5068
− .3665
− .5365
− .1272
− .3771
Appendix
Table 6: Volatility IC by resolution. Higher is better; bold marks the best value in each column.
Configuration
IC/RankIC ↑
KiT
0.0372/0.0407
Kronos-base
0.0090/0.0148
Kronos-base-FT
0.0132/0.0205
Appendix
Table 7: Price-series forecasting. Values are averaged over the seven resolutions and three markets.
Identity wid ( whist=1 )
History whist ( wid=1 )
w
Prior ↓
Hist. ↓
Spread ↓
Prior ↓
Hist. ↓
Spread ↓
1
0.41
0.52
1.08
0.41
0.52
1.08
2
0.28
0.49
0.91
0.39
0.37
0.93
3
0.20
0.47
0.78
0.37
0.26
0.81
4
0.15
0.46
0.69
0.36
0.19
0.72
Appendix
Table 8: Guidance sweep averaged over 100 products.
Representation of a bar
Mean return RankIC ↑
Five-dimensional log-ratio state ( KiT )
0.0571
Close-only log-price input
0.0183
Naive OHLCV input (normalized)
0.0209
Appendix
Table 9: Ablation on the candlestick representation. Each row alters only the input state; all other settings follow the main model.
Configuration
Mean return RankIC ↑
AdaLN identity + per-bar calendar ( KiT )
0.0571
Identity → AdaLN
Market and sector only (drop instrument, scale)
0.0386
w/o identity
0.0329
Per-bar condition → token embeddings
w/o per-bar condition
0.0372
Appendix
Table 10: Ablation on conditioning injection. Identity denotes the window-constant signals (market, sector, instrument, timescale); Calendar denotes the per-bar signals (clock, session, day of week, month, day of year, event flags).
This technical report presents EXAONE Forecast for Finance (EXAONE Finance), a financial time series (TS) foundation model (TSFM) tailored to financial forecasting. Recent TSFMs achieve strong zero-shot performance through large-scale pretraining. However, they are primarily developed for general-domain TS and largely rely on self-attention backbones whose computational cost grows quadratically with sequence length and variate count. Moreover, they assume fully observed inputs and are pretrained on corpora that fail to capture the unique dynamics of financial markets. These limitations hinder their applicability to finance, where long, many-channel, intermittently observed panels are common. To address these challenges, EXAONE Finance adopts an attention-free architecture, replacing self-attention with two simple yet effective linear-time operators: 1) a causal 1D convolution for temporal mixing and 2) a group-aware pooling multi-layer perceptron (MLP) for variate mixing. Furthermore, a masked context augmentation exposes the model to contiguous missing spans during training, improving robustness to the missingness pervasive in financial markets. EXAONE Finance is pretrained on a large-scale financial corpus covering not only equities but also foreign exchange, commodities, crypto-assets, fixed income, and macroeconomic indicators. On FinVerse, a financial forecasting benchmark covering diverse asset classes, EXAONE Finance attains state-of-the-art performance, ranking first across all three evaluation tiers---point-forecast accuracy, cross-sectional asset ranking, and portfolio profitability.
While Next-Token Prediction (NTP) has unified LLM pretraining, its adaptation to unbounded, continuous time series (TS) remains open. To bridge the gap, we introduce UniTok, a universal tokenizer that transforms TS into discrete tokens, and UniTok-FM, a foundation model pretrained via NTP on these tokens. UniTok-FM is a general-purpose foundation model that supports zero-shot and prompt-boosted forecasting, as well as few-shot generation and classification via training-free in-context inference--a capability not achieved by prior works. Technically, UniTok is a vector-quantized autoencoder incorporating prefix normalization for scale stabilization, a progressive-resolution causal architecture for encoding and decoding, and a structure-preserving reconstruction loss for training. UniTok-FM adopts an off-the-shelf LLM architecture without TS-specific modifications. Instead of pretraining on isolated TS, it performs NTP on context windows formed by multiple series with similar patterns, aiming to capture their shared dynamics. Experiments on forecasting, generation, and classification show that a single unified UniTok-FM consistently outperforms statistical and supervised baselines, achieves competitive performance with task-specific foundation models, and uniquely enables training-free in-context inference across tasks.
Yunhao Zhang, Ruiying Qi, Jiale Zheng +3
Shanghai Jiao Tong University · Huawei Noah’s Ark Lab
Probabilistic K-line forecasting describes uncertainty in four complementary prices, namely open--high--low--close (OHLC). However, it introduces two consistency problems: quantile crossing and K-line crossing. Quantile crossing occurs when a higher-quantile forecast falls below a lower-quantile forecast, while K-line crossing occurs when the forecast low exceeds the open or close, or the forecast high falls below the open or close. Existing solutions generally address only one problem through output reordering, specialized architectures, or penalized training objectives. We propose K-line--Quantile Sequential Projection (KQSP), a parameter-free and training-free reconciliation method applicable to forecasts produced by any model. Compared with other crossing solutions, KQSP preserves predictive accuracy while producing substantially smaller corrections to the original forecasts. To mitigate model bias, we evaluate KQSP using various models, including pretrained foundation models. KQSP reduces both quantile and K-line crossing rates to zero for all test data undertaken. These results show that probabilistic K-line consistency can be enforced independently of forecast generation and without retraining.
Runyao Yu, Yuchen Tao, Yujie Chen +2
London Business School London, United Kingdom · RWTH Aachen University Aachen, Germany · The Chinese University of Hong Kong Shenzhen, China +1