EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
Authors: Takumi Fujimoto, Hiroaki Nishi
Organizations: School of Science for Open and Environmental Systems, Graduate School of Science and Technology, Keio University, Yokohama 223-8522, Japan · Department of System Design, Faculty of Science and Technology, Keio University, Yokohama 223-8522, Japan
Completed forecasts provide residual feedback, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Each channel shares the endpoint across component-wise online ridge regressions using current-forecast coefficients. We evaluate eight multivariate series, DLinear and PatchTST, three seeds, and two training variants: 96 matched fixed-base conditions at a 24-step horizon. EPOC achieves mean condition-wise reductions in mean squared error (MSE) and mean absolute error (MAE) of 15.40% and 9.35% from the uncorrected base, respectively, with a median of 6,352 B in retained auxiliary arrays. EPOC also outperforms the δ-Adapter, COSA, FAC, and OMPB in paired MSE on most conditions while retaining less state. Full ELF achieves the largest mean MSE reduction, 19.29%, but its median retained state is 474,048 B (×75 relative to EPOC). Equal-size summary controls favor the endpoint by 1.65--2.20% in paired MSE; a coefficient-reconstructed endpoint yields similar accuracy to the observed endpoint, highlighting its role as a shared input. Increasing the retained DCT component count from 4 to 8 adds 1.00 percentage point of MSE reduction for 5,728 B. On jointly trained bases, EPOC lowers MSE by 16.69--20.15% relative to globally blended TEFL-style adapters applied to the same base. Code and numerical records are available at https://github.com/keiotakmin/endpoint-preserving-residual-correction.
Figures & tables
Method
Auxiliary input
Correction and feedback
Retained state
ELF (ICML 2025)
Fourier features of observed history
Online auxiliary forecast and adaptive mixing
Regression and mixing statistics
TEFL (arXiv 2026)
Most recent completed residual block
Low-rank residual adapter; joint base training
Adapter weights and residual input
δ -Adapter (ICLR 2026)
Current input and forecast
Bounded input/output adapters; online updates
Adapter weights and optimizer state
COSA (ICLR 2026)
Current forecast and recent target statistics
Gated linear output correction; online updates
Linear adapter, gate, context
FAC (arXiv 2026)
Frequency representations of input and forecast
Gated frequency-domain calibration
Masks, biases, optimizer state
OMPB (UAI 2026)
Base forecast
Gated Bayesian residual head; predict then update
Head, optimizer, source/target buffers
Table 1: Auxiliary inputs, correction mechanisms, and retained state in the evaluated comparison methods. Existing methods are ordered by publication year; EPOC appears last.
Figure 1: Correction of block t and the complete-block feedback order. The DCT coefficients and final value of et−1 are available when ft is issued. The residual trace is schematic. After yt is fully observed, its base residual updates the regression, the blending controller, and the summary used at the next origin.
Series
Samples
C
Eval.
Blocks
ETTh1
17420
7
8710
362
ETTh2
17420
7
8710
362
ETTm1
69680
7
34840
1451
ETTm2
69680
7
34840
1451
Appliances
19735
26
9867
411
BDG2 (Rat)
17544
15
8772
365
Table 2: Prepared evaluation series. Eval. is the zero-based index of the first evaluation target; Blocks counts complete 24-sample targets. All eight series enter the same fixed-base comparison.
96-condition overview
Method
MSE ↑ (%)
MAE ↑ (%)
Retained bytes ↓
Static
0.00
0.00
0 ( ×0 )
δ -Adapter
6.70
5.12
22,552,608 ( ×3550 )
COSA
11.73
7.56
111,756 ( ×18 )
FAC
14.59
9.11
33,312 ( ×5 )
OMPB
10.05
6.73
767,216 ( ×121 )
Table 3: Seven-method comparison on 96 matched conditions. The overview averages condition-wise reductions from Static; retained-byte multiples are relative to EPOC (6,352 B) and rounded to integers. The MSE and MAE panels average losses over three seeds and the legacy/refit variants within each series–backbone pair. The rightmost ratio divides EPOC’s unrounded loss by the lowest unrounded loss in that row; 1 indicates the best result. Bold marks row-minimum losses and the corresponding ratio when EPOC is best. All corrections use global blending.
Figure 2: Accuracy and retained correction state by channel group for six methods on 96 fixed-base conditions. Left bars show mean Static-relative MSE reduction: each series–backbone value averages three seeds and two training variants, and bars then average four ETT series, three BDG2 sites, or one Appliances series. Hatched bars denote DLinear and solid bars PatchTST. Whiskers show the range across series, not statistical uncertainty; Appliances has one series and no whiskers. Right bars show retained numerical-array bytes for each method and group on a logarithmic scale; this value is common to both backbones and all series in the group.
Same-base: EPOC vs. adapter
Separate training: adapter vs. base-only
Retained state (B)
Warmup
Width
MSE (%)
MAE (%)
MSE (%)
MAE (%)
EPOC
Adapter
MAE
2
19.27
10.78
2.22
1.29
4,528
1,584
4
18.98
10.61
3.13
1.77
4,528
1,968
64
16.69
8.82
6.09
3.68
4,528
13,488
MAE+SF
2
20.15
11.29
2.49
1.27
4,528
1,584
4
19.54
10.86
3.68
1.98
4,528
1,968
Table 4: Joint-training comparison over six series, two backbones, and three seeds. The same-base columns give the paired reduction of EPOC relative to the globally blended adapter on the same jointly trained base. The separate-training columns give the reduction of the raw joint adapter relative to a separately trained base-only run. For each metric, seed losses are averaged within each of 12 series–backbone pairs before calculating reductions, which are then averaged across pairs. Retained bytes are medians for the globally blended corrections. SF denotes spectral flatness in base warmup; positive reductions favor the first-named method.
MSE reduction (%)
MAE reduction (%)
Base / training
First
Middle
Mean
First
Middle
Mean
Retained (B)
Frozen base (96 conditions)
2.20
1.65
1.96
1.46
1.11
1.39
6,352
Joint MAE (12 pairs)
4.70
4.83
6.00
2.93
2.95
3.61
4,528
Joint MAE+SF (12 pairs)
5.43
5.49
6.94
3.38
3.34
4.14
4,528
Table 5: Endpoint gain over equal-size residual summaries. Each entry gives the mean paired reduction when the preceding block’s endpoint replaces the listed first residual, middle residual, or block mean. The frozen-base row averages 96 condition-wise reductions; each joint row averages reductions across 12 series–backbone pairs after seed-mean losses are formed. All variants within a comparison use the same base, regression settings, blending rule, and retained array size.
Figure 3: Input construction on 96 matched fixed-base conditions. Markers show mean condition-wise MSE and MAE reductions from Static; bars show median retained bytes with global blending. The projected endpoint r~ is reconstructed from retained residual DCT coefficients and supplied to every component regression. Endpoint persistence repeats the preceding endpoint without a fitted regression.
Figure 4: Main comparison and one-factor sensitivity on 96 fixed-base conditions. Upper panels plot mean condition-wise reductions from Static against median retained bytes for the six globally blended correction methods. Lower left varies the retained DCT component count K ; lower right changes ridge strength, forgetting half-life, or blending window relative to the K=4 reference. Unchanged settings follow Section 3 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Setting
Lookback / horizon / stride
96 / 24 / 24
EPOC residual basis
DCT, K=4
Ridge / forgetting
λ=1 , ρ=2−1/128
Blending window
32 completed blocks
Learned widths / seeds
2, 4, 64 / 0, 1, 2
Warmup / joint training
3 / 12 epochs
Appendix
Table S1: Common evaluation settings and learned-adapter training configuration for the supplementary records. The fixed-base training and comparator protocols are given in Section 4 .
Setting
MSE (%)
MAE (%)
Bytes
Reference ( K=4 )
15.40
9.35
6,352
K=1
11.13
6.47
2,056
K=2
13.10
7.85
3,488
K=3
14.49
8.79
4,920
K=6
16.00
9.76
9,216
K=8
16.40
10.08
12,080
Appendix
Table S2: One-factor sensitivity on 96 fixed-base conditions. Percentages average condition-wise reductions from Static with global blending; bytes are median retained arrays. Other settings match the reference.