Prediction Limits and Koopman Closure of Geometry-Induced Soft State Abstractions
Organizations: Software Competence Center Hagenberg, Hagenberg, Upper Austria, Austria
Abstract
A soft state representation assigns each state a vector of nonnegative class weights that sum to one. We study how the construction of these weights and the state dynamics jointly determine the accuracy of linear prediction. For any fixed measurable representation, we derive a finite-sample lower confidence bound on the smallest population root-mean-square prediction error among matrices with a specified spectral-norm limit. The bound compares variation in successor coordinates within each reference class with the improvement that soft inputs could provide. It is computed from independent evaluation pairs without fitting a prediction matrix. A bound above a chosen tolerance rules out that tolerance for the entire matrix class; a zero bound is inconclusive. For coordinates constructed using Kernel Affine Hull Machines, reconstruction-score margins control disagreement with reference labels and enter bounds on prediction error. Under exact deterministic linear evolution, we also establish the Koopman and reproducing-kernel Hilbert-space adjoint interpretation, accounting for redundant coefficient vectors. A four-state study compares the confidence bound with analytically known optima across 117,000 reported replicate datasets. A Van der Pol representation selected on pilot data is then evaluated on 32 independent datasets under each of two transition laws. The reported bounds are positive at the fitted matrix norm, but can become zero at larger norm limits. Further forecasting studies examine coordinate variation, common prediction targets, and long-horizon error. The results distinguish agreement with reconstruction classes, attainable prediction accuracy, and exact operator closure.
Figures & tables
| Question | Supporting result | Conditions and interpretation |
|---|---|---|
| Can any permitted linear predictor meet a chosen error tolerance? | Population lower bound and finite-sample certificate. | Fixed representation and evaluation law; a bound above the tolerance excludes it for the norm-bounded class. A zero bound is inconclusive. |
| How does reconstruction quality affect prediction? | Bounds based on reconstruction-score margins and a four-state example. | The margin bounds control agreement with labels. Successor variation must also be considered. |
| When is coordinate prediction an exact Koopman representation? | Finite-feature RKHS and reduced spectral correspondence. | Deterministic dynamics, exact closure on the domain, and nonzero represented functions. |
| How well do fitted models forecast the chosen coordinates? | Forecasting experiments on specified targets. | Conclusions depend on the target, horizon, selection procedure, and sampling unit. |
| Symbol | Meaning |
|---|---|
| , , | Current soft coordinates, successor soft coordinates, and the current one-hot reference label. |
| Minimum population root-mean-square error over matrices with . | |
| Root-mean-square variation of successor coordinates about their class mean. | |
| Root-mean-square Euclidean distance between the soft coordinates and the one-hot label. | |
| Mean squared shortfall of the weight assigned to the reference class; . | |
| , | Evaluation-sample assignment loss and within-class successor variance. |
| Bound from uniform control of risk | Bound based on reference labels | |
|---|---|---|
| Empirical quantity | Minimum sample prediction error over the matrix class, or a certified lower bound on it. | Assignment loss and within-class successor variance. |
| Role of labels | Uses coordinate inputs and successor targets directly. | Groups successor vectors by their current reference label. |
| Use of soft coordinates | Finds their best linear predictor within the matrix class. | Bounds the possible reduction in error relative to hard-label prediction. |
| Main limitation | Needs a uniform confidence correction and a valid lower bound on minimum empirical risk. | May be zero despite positive minimum error, especially with few samples per class or large deviations from one-hot labels. |
| System | Evaluation law | Critical norm | |||
|---|---|---|---|---|---|
| Duffing | Matched stochastic | 1.005768 | 0.780756 | 0.000217 | 0.011796 |
| Duffing | Deterministic | 1.005768 | 0.780787 | 0.000222 | 0.011914 |
| Van der Pol | Matched stochastic | 1.078133 | 0.360983 | 0.017542 | 0.155876 |
| Van der Pol | Deterministic | 1.078133 | 0.361292 | 0.017328 | 0.154858 |
| Evaluation law | Positive | Median [5th, 95th] | Minimum | Median RMSE |
|---|---|---|---|---|
| Matched stochastic | 32/32 | 0.084727 [0.078444, 0.089910] | 0.076333 | 0.276233 |
| Deterministic | 32/32 | 0.084269 [0.077720, 0.092126] | 0.075401 | 0.275797 |
| Fitted matrix | Positive | Median | |||
| Norm limit | permitted? | Stochastic | Deterministic | Stochastic | Deterministic |
| No | 32/32 | 32/32 | 0.089186 | 0.088660 | |
| Yes | 32/32 | 32/32 | 0.084727 | 0.084269 | |
| Yes | 11/32 | 8/32 | 0 | 0 | |
| Yes | 0/32 | 0/32 | 0 | 0 | |
| Dynamics | Optimal RMSE | Certificate limit | |
|---|---|---|---|
| Identity: | |||
| Conflicting successors: ; |
| Mean | Centered error | Selected | |||||
|---|---|---|---|---|---|---|---|
| 25 | 6 | 0.8885 | 0.3545 | 14.26 | 0.5942 | Yes | |
| 30 | 6 | 0.8865 | 0.3219 | 15.58 | 0.5373 | No | |
| 20 | 4 | 0.8865 | 0.2720 | 10.42 | 0.5482 | No |
| Duffing, , | Van der Pol, , | |||||
|---|---|---|---|---|---|---|
| Coordinate map | ||||||
| KAHM affinities | 0.9942 | 0.9976 | 0.180 | |||
| -means RBF | 0.9950 | 0.9672 | 0.819 | |||
| -means distance | 0.9924 | 0.583 | 0.9001 | 1.07 | ||
| -means hard | 0.9411 | 0.393 | 0.112 | 0.8792 | 1.09 | |
| Training trajectories | Across noise levels | No added noise | Reference | ||
|---|---|---|---|---|---|
| 1 | |||||
| 2 | |||||
| 3 | |||||
| 5 | |||||
| 8 |
| System | Method | |||
|---|---|---|---|---|
| Duffing | Coordinate NLMS | |||
| State DMD | ||||
| State EDMD, degree 2 | ||||
| State EDMD, degree 3 | ||||
| Van der Pol | Coordinate NLMS | |||
| State DMD |
| Task | |||||
|---|---|---|---|---|---|
| CartPole | |||||
| MountainCar | |||||
| Acrobot |
| Predictor | ||||
|---|---|---|---|---|
| Direct coordinates, NLMS | ||||
| Affine state predictor | ||||
| Quadratic state predictor | ||||
| Cubic state predictor |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Study family | Role here | Selection rationale |
|---|---|---|
| Four-state certificate evaluation (21) | Checks against an analytical optimum. | The known population optimum allows coverage checks. Counts and tolerance decisions are reported separately from the dynamical pilot. |
| Dynamical pilot and subsequent certificate evaluation (22-23) | Application to one selected representation. | Zero pilot outcomes, evaluation on new pairs after selection, and results at all four norm limits are included. |
| Coordinate construction (6-7) and selection by centered error (9, 12) | Representation diagnostics and hyperparameter selection. | Each map predicts its own coordinates; variation diagnostics help interpret these changing targets. |
| Increasing training-trajectory count (13) | Effect of additional training data. | All five training counts and three configurations are included. Each map is reconstructed and predicts its own coordinates. |
| State-space predictor baselines (15) | Comparison on a common target. | Direct coordinate and state-space predictors share an output map within each oscillator system. Both systems and all reported methods are included. |
| Classic Control and Acrobot (16, 18) | Main fixed-policy diagnostics. | One-step and 50-step scores are both retained, including negative long-horizon outcomes. The measured outcome is coordinate prediction under fixed policies. |
| Study | Count | Objective | ||
|---|---|---|---|---|
| Duffing reference | 9 | |||
| Duffing expanded | 64 | |||
| Van der Pol, no added observation noise | 49 | ; 1-SE | ||
| Van der Pol, across observation-noise levels | 30 | weighted ; 1-SE | ||
| Acrobot | 30 |
| Operation | Git revision prefix |
|---|---|
| Fix confirmatory protocol and configuration | 18c2bb4 |
| Build training-only representation and matrix | 4a0cedd |
| Recompute the representation and matrix for verification | cbfe438 |
| Generate 64 independent-pair evaluation datasets | 18dc9f3 |
| Recompute results from the completed evaluation | 090eac5 |
| Generate summaries and repeat numerical verification | 689d40d |
| System | Rollout rule | |||
|---|---|---|---|---|
| Duffing | No projection | |||
| Per-step vector projection | ||||
| Row-stochastic matrix | ||||
| Van der Pol | No projection | |||
| Per-step vector projection | ||||
| Row-stochastic matrix |
| System | Estimator | Min. | Median | Max. | Counts | ||||
|---|---|---|---|---|---|---|---|---|---|
| Duffing | NLMS | 0.9835 | 1.000 | ||||||
| Ridge least squares | 0.9842 | 1.000 | |||||||
| Van der Pol | NLMS | 0.9963 | 1.000 | ||||||
| Ridge least squares | 0.9964 | 1.000 |