Learning-to-Defer in Non-Stationary Time Series via Switching State-Space Models
Organizations: School of Computing National University of Singapore · Université de Toulouse Fédération ENAC ISAE-SUPAERO ONERA · Institute for Infocomm Research A*STAR, Singapore
Abstract
Learning-to-defer (L2D) lets a predictor decide, at each round, whether to issue its own forecast or pay for an expert's. In non-stationary time series this decision must keep adapting, although deployment reveals only the consulted expert's forecast while a historical archive records every expert with the target. L2D-SLDS learns from this archive a switching state-space model of the target and all expert forecasts, whose shared and expert-specific states describe how experts move together and apart. Its predictive law supplies the internal forecast and the expected cost of every consultation, which a greedy router minimizes, and one consultation also updates the beliefs about unconsulted and unavailable experts. We prove sublinear regret against a changing conditional-risk oracle without exploration, when the candidate models are accurate and either the archive separates them or live feedback reveals cost differences. On three real datasets, L2D-SLDS has the lowest cost among eight bandit routers and adapts its consultation rate to the fee.
Figures & tables
| System | Delhi | Jena | Melbourne |
|---|---|---|---|
| L2D-SLDS | |||
| LinUCB | |||
| Shared LinUCB | |||
| Linear TS | |||
| Ensemble | |||
| Discounted LinUCB |
| Variant | Delhi | Jena | Melbourne |
|---|---|---|---|
| Greedy (ours) | |||
| Inverse-gap | |||
| Joint IX | |||
| Structured | |||
| Retain | |||
| Reset |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Number of external experts, regimes, and predictive candidates | |
| Available external experts; | |
| Selected action, internal forecast, expert forecast, target | |
| Output bound, access fee, maximum action cost | |
| Observed pre-action information | |
| Shared state, private state, regime |
| Dataset | Prefix | Fit | Warmup | Validation | Heldout | Heldout target dates |
|---|---|---|---|---|---|---|
| Delhi | 627 | 313 | 78 | 78 | 472 | 2016-01-02–2017-04-24 |
| Jena | 2400 | 1200 | 300 | 300 | 1800 | 2010-12-02–2011-09-27 |
| Melbourne | 1458 | 729 | 182 | 182 | 1096 | 1987-12-31–1990-12-31 |
| Method | Tuned settings and fixed settings |
|---|---|
| LinUCB / Shared LinUCB | Bonus multiplier ; ridge . |
| Linear TS | Posterior scale ; ridge . |
| Ensemble | Perturbation standard deviation ; ridge ; 20 members. |
| Discounted LinUCB | Discount ; bonus multiplier ; ridge 1, noise scale , parameter-norm bound 1, confidence parameter . |
| CUSUM-LinUCB | Bonus ; threshold ; slack ; warmup and reference length 25, ridge 1. |
| GLR-LinUCB | Bonus ; detector level ; minimum window ; reference length 25, ridge 1. |
| Fee | L2D-SLDS cost | Best contextual baseline | Baseline cost | L2D-SLDS queries (%) |
|---|---|---|---|---|
| 0 | Ensemble | |||
| 0.05 | Ensemble | |||
| 0.10 | LinUCB | |||
| 0.20 | LinUCB | |||
| 0.40 | Shared LinUCB | |||
| 0.80 | LinUCB |
| Dataset | Policy | Prediction loss | Fee | Query fraction | Internal MSE |
|---|---|---|---|---|---|
| Delhi | L2D-SLDS | ||||
| Inverse-gap | |||||
| Joint IX | |||||
| Jena | L2D-SLDS | ||||
| Inverse-gap | |||||
| Joint IX |
| Data | Full | One regime | |
|---|---|---|---|
| Switching | |||
| Stationary |
| Dataset | L2D-SLDS | LinUCB | Disc. LinUCB | Linear TS | Reset | |
|---|---|---|---|---|---|---|
| Delhi | ||||||
| Jena | ||||||
| Melbourne |
| Policy | Delhi | Jena | Melbourne |
|---|---|---|---|
| L2D-SLDS (structured) | |||
| Inverse-gap | |||
| Joint IX |
| Prediction cost | Posterior (KiB) | ||||
|---|---|---|---|---|---|
| Full | Structured | Full | Structured | Reduction | |
| 4 | 0.727 | 0.352 | 51.6% | ||
| 8 | 2.133 | 0.633 | 70.3% | ||
| 16 | 7.195 | 1.195 | 83.4% | ||
| 32 | 26.320 | 2.320 | 91.2% | ||
| 64 | 100.570 | 4.570 | 95.5% | ||
| Method | |||||
|---|---|---|---|---|---|
| L2D-SLDS (full) | |||||
| L2D-SLDS (structured) | |||||
| LinUCB | |||||
| Shared LinUCB | |||||
| Linear TS | |||||
| Ensemble |
| Method | Posterior (KiB) | Policy buffers (KiB) | Policy time (ms) | Peak RSS (MiB) |
|---|---|---|---|---|
| L2D-SLDS (full) | 100.570 | 0.000 | ||
| L2D-SLDS (structured) | 4.570 | 0.000 | ||
| LinUCB | 100.570 | 6.602 | ||
| Shared LinUCB | 100.570 | 542.945 | ||
| Linear TS | 100.570 | 542.945 | ||
| Ensemble | 100.570 | 666.227 |