Decision-focused learning (DFL) trains predictors through downstream objectives, but a different loss need not provide an independent parameter-update direction. We characterize this restriction through the predictor Jacobian, using sparse index tracking to distinguish the covariance entries read by the optimizer from the parameter directions available to learning. Rank-one Jacobians make nonzero per-example gradients collinear; a conditional spectral bound describes near-collinearity. A batch-subspace characterization and counterexamples show why these local statements imply neither common minimizers nor collinear batch updates. Experiments examine when geometry translates into decision quality. Across 38 one-parameter equity configurations, DFL gains over MSE remain below 1.8%; a 385-parameter conditional predictor also has pointwise rank one. In validation-tuned shortest-path and knapsack experiments, full-capacity SPO+ reduces mean regret by 11.6% and 10.6%, respectively; only knapsack survives correction across eight comparisons. The capacity contrast persists on fresh datasets across batch orders and training budgets. Holding expressivity fixed, invertible coordinate scaling lowers spectral effective rank and ordinary SGD gains; compensating for the scaling restores the original trajectories. Financial forward-target controls separate forecast accuracy from decision quality; a matched neural comparison finds no aggregate DFL advantage in the tested architecture. These findings distinguish local rank restrictions, coordinate-dependent optimization and predictive accuracy. Predictor geometry helps explain available learning directions, while held-out decision quality remains the test of practical benefit.
Figures & tables
Figure 1 : Local geometry restricts directions; task gains require evidence. Left: rank-one parameter Jacobians map nonzero loss gradients onto one line, but batch directions can differ across examples. Right: full-capacity regret reductions in the controlled extension, with pointwise 95% paired-dataset bootstrap intervals over ten datasets. Only knapsack passes Holm correction across eight comparisons. Filled markers denote Holm-adjusted significance.
Figure 2 : Two objectives, one predictor. MSE directly supervises covariance prediction; task loss differentiates through the QP. Both reach θ through J⊤ , with selection detached. The toy example illustrates sparse weighting for an equal-weight index with identity covariance.
Figure 3 : Where gradient directions are lost. (a) The task-relevant support has 2NK−K2 entries (36% here). (b) A rank-one predictor Jacobian maps nonzero gradients onto one line, with either sign. (c) With multiple active singular directions, angular separation can survive; improved task performance is possible, not guaranteed.
Figure 4 : Pointwise Jacobian measurements on 54 inputs per model class. A 385-parameter conditional predictor is numerically rank one because its output passes through a scalar shrinkage intensity. Its gradient direction can still vary across inputs. These measurements distinguish parameter count from local rank; they do not measure the stacked batch rank.
Figure 5 : Validation-tuned cross-domain extension. (a) Mean regret reduction with pointwise 95% paired-dataset bootstrap intervals (ten datasets; 10,000 resamples). (b) Mean absolute MSE/SPO+ batch-gradient cosine at the selected MSE model, using 32 held-out examples. Each capacity is tuned separately; full capacity has 144/72 update directions for path/knapsack. Only full-capacity knapsack passes Holm correction across eight tests.
Figure 6 : Capacity contrast on fresh data and longer training. Mean test-regret reduction versus validation-selected MSE; whiskers are pointwise 95% paired-dataset bootstrap intervals. Each task uses ten new datasets and three minibatch orders, averaged within dataset before inference. Both losses may select the unchanged initializer. Filled markers denote comparisons passing Holm correction across eight tests; only full-capacity knapsack passes. Scalar gains below 0.6% are not evidence of equivalence.
Figure 7 : Same function class, different ordinary SGD behavior. Invertible coordinate scaling reduces spectral effective rank while preserving exact rank and expressivity. Compensated SGD recovers the original predictor updates. Whiskers are pointwise 95% paired-dataset bootstrap intervals over ten fresh datasets per task, not multiplicity-adjusted tests.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Loss choice, model capacity, and selection stability. (a)–(b) Tracking error relative to POET; (c) relative to MSE (lower is better). At N=100 , the three fitted one-parameter estimators improve on Ledoit–Wolf but meet the empirical equivalence criterion in Table 1 . At N=478 , estimator choice changes the ordering. Under dynamic selection, BD-DFL reduces the extreme errors of unregularized neural DFL; dots show individual folds.
Figure 9 : Tracking-error reduction over nine paired folds. Bars show 100(1−TEDFL/TEbase) ; whiskers are ±1 paired delta-method SE of this ratio, a descriptive summary across folds. Unadjusted one-sided Wilcoxon: ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 . (a) Neural DFL vs. MSE: reductions of 7.0–16.5% at K≤30 , and 0.2% at K=50 . (b) Shrinkage DFL vs. Ledoit–Wolf: reductions of 7.1–19.1%; comparisons with other fitted one-parameter estimators are in Table 1 .
Figure 10 : Corrected rolling tracking error for the neural model at K=20 : 2,096 observations across nine test folds. Top: annualized 63-day standard deviation of tracking returns. Bottom: relative TE reduction from DFL. Dotted lines mark fold boundaries; gaps exclude each fold’s first 62 observations.
Figure 11 : Task training across factor counts and portfolio cardinalities (nine paired folds per cell). Values are relative TE changes, 100(TEDFL/TEMSE−1) ; negative values indicate improvement. In this grid, changes vary more across K than across Kf , and improvements are smaller at K=50 . Unadjusted one-sided Wilcoxon: ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 . The heatmap does not identify a causal effect of either axis.
Figure 12 : DFL gains for three model classes at N=100 , using nine paired folds. Baselines are validation-tuned shrinkage (one parameter) and MSE training (structured and conditional models). Bars show relative reduction in mean TE; whiskers are ±1 paired delta-method SE. Unadjusted one-sided Wilcoxon: ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 . The 12-parameter model has the largest observed reduction here; the ordering differs at N=478 (Table 2 ). Architecture changes alongside parameter count.
Figure 13 : Regularization under dynamic selection. (a) Shrinkage at K=20 across the tested coefficients; dashed segments bridge untested settings. (b) Best-tested coefficient for each model and sparsity at N=478 . These exploratory comparisons do not establish that regularization gains grow monotonically with universe size.
Figure 14 : Tracking-error–turnover trade-off for the tested neural configurations at K=20 . Points are fold means; connectors join the sampled regularization coefficients and do not certify a continuous Pareto frontier. The MSE reference is shown separately. At γ=0.01 , mean TE is lower than MSE while turnover is higher (0.005 vs. 0.003; Table 15 ); the preferred trade-off depends on trading costs.
Figure 15 : Archived blind-evaluation record (33 matched cases). Colour shows DFL gain: red >1% , grey within ±1% , blue <−1% . (a) Four incorrect predictions are ringed; overlapping finance points are counted. (b) Equal spacing denotes the four sampled categories, not a continuous sweep; horizontal marks show means. This historical snapshot includes two abstentions.
Figure 16 : Synthetic heterogeneity sweep at N=100 , K=20 : 18 settings, five folds each. Faint dots show every fold; larger markers show mean per-fold TE reduction and ±1 SEM. The ±1.6% band is a descriptive reference, not a confidence interval; individual folds can lie outside it. Mean gains are larger at several settings above the empirical split at dproxy=9.5 , but fall to approximately 1.0% at the final setting. The split is not a validated universal threshold.
Figure 17 : Shortest-path experiment on an 8×8 grid with five paired seeds. (a) Mean SPO+ advantage in percentage points, Δ=regretMSE−regretSPO+ , with ±1 SEM. Bottleneck width k bounds input-Jacobian rank; full linear and MLP models also change architecture. The low-rank point has a small negative estimate, while intermediate ranks are mixed. (b) Mean normalized regret for the same paired runs; labels give the SPO+ change in percentage points. Shading describes sampled rank ranges, not a theorem about test regret.
Figure 18 : SPO+ advantage in percentage points at two grid sizes; positive values favor SPO+. Whiskers show ±1 SEM across five paired seeds. Filled markers join the controlled bottleneck settings; open diamonds denote full-capacity architectures. The larger grid has greater observed advantage at the sampled settings. A SEM bar crossing zero is not a hypothesis test; paired inferential results are reported separately in the corresponding table.
Figure 19 : Two empirical axes, with a shared vertical scale. (a) Thirty-eight one-parameter equity configurations across six markets and a 54× range of K/N : gains remain below 1.8%, with no detected monotone association ( ρ=−0.16 , p=0.35 ). The shaded ±1% band is a reference, not an envelope containing all points. (b) PyEPO capacity sweep: relative reduction in mean regret, with ±1 paired delta-method SE across ten seeds. Dotted thresholds were calibrated on equities; one transfer failure is marked. The two panels concern different tasks and support association rather than a causal comparison.
Figure 20 : Improved covariance fit need not improve tracking. Left: annual held-out TE reveals the uneven cost of future-target fitting across market years. Right: changes in two distinct metrics relative to reconstruction; negative is better for either metric. Their percentages have different denominators and are not a common utility scale.
Figure 21 : A matched neural control without an aggregate DFL gain. (a) Bars compare seed-averaged TE within each year; dots show the three paired initialization results and are not additional independent datasets. Positive values favor DFL. (b) Validation-selected checkpoints across 30 models per loss. Frequent no-update selections limit conclusions about trained-model superiority.
Department of Industrial Engineering and Operations Research University of California, Berkeley · H. Milton Stewart School of Industrial and Systems Engineering Georgia Institute of Technology