The Arbitrary-Placement Problem in Entropy-Minimizing Selection, and a Residual-Entropy Formulation
Authors: Alyssa H. Shin, Claire H. Shin
Organizations: Division of Biology and Biological Engineering California Institute of Technology · Robert Frederick Smith School of Chemical and Biomolecular Engineering Cornell University
Entropy-based selection objectives suffer from a fundamental degeneracy: minimizing Shannon entropy H(pA) rewards confident selection regardless of whether the selected candidate is informative. We address this limitation with the residual entropy D=H(pA)−H(pβ), where pβ is induced by candidate trust weights. We prove the exact identity D=−KL(pA∥pβ)−Δ, where Δ measures whether the score-induced distribution and trust profile favor the same candidates. Boundary cases establish basic safety: under uniform trust, D≤0 automatically, so an equal-trust, non-starving state is never penalized, while at any one-hot limit, D→0 regardless of the selected candidate. For the intermediate regime where selection occurs, we prove that D≤0 when candidate ordering by trust agrees pairwise with ordering by informativeness, and derive a tighter certificate based on the leading candidate's margin over its competitors. These results are independent of the candidate-scoring function and apply to both stationary and dynamically changing information. Experiments with a gradient-based mixture-of-experts router confirm that the ordering conditions can hold during real optimization and show that correct ordering improves downstream performance when candidates are non-interchangeable and selections are used directly rather than averaged. Beyond routing, margin-based reweighting matches or outperforms fixed-strength baselines in a class-imbalance task, while informative selection in a production video-prediction system reduces MSE by approximately 20% and transfers to a related species. Residual entropy, therefore, provides a safety criterion for selection and a usable signal for deciding when that selection is informative.
Figures & tables
Figure 1: From entropy degeneracy to residual-entropy correction. Propositions 1–2 characterize the problem and correction and Corollaries 1–4 provide boundary results and sufficient sign certificates.
−0.3641 / −0.0134 ; lower objective at informative initialization
(iv) Class imbalance: identification
Selection rate 0.775
0.938 ( 1.21× ; p=0.003 )
(v) Full-batch MoE: participation boundary
5/6 lower-bounded at floor >0 ; no transition
Threshold 0.07–0.10; 0/6 below, 4/6 above
(vi) Streaming MoE: boundary robustness
5/6 lower-bounded at floor >0 ; no transition
Threshold 0.062–0.070; 0/6 below, 4/6 above
Table 1: Seven studies: questions and quantitative evidence. Objective pairs are informative / elsewhere initialization means.
Setting
Uniform
Fixed 5×
Conf.-scaled
Hurt (fixed / conf.)
p (conf. vs. fixed)
Digits T=5 , 8% ( n=80 )
0.8466
0.8727
0.8732
2/80 / 0/80
0.159
Digits T=5 , 2% ( n=15 )
0.6776
0.6909
0.7018
5/15 / 0/15
0.023
Table 2: Confidence-modulated reweighting at both scarcity levels: mean held-out minority-class recall after retraining with each oversampling scheme. “Hurt” counts seeds below uniform training; the final column compares confidence-scaled with fixed 5× weighting. The same s=0.25 is used.
Noise
Threshold
Clamped acc.
Peaked acc.
Winner
0.0
0.10–0.12
0.960
0.964
peaked ( +0.3 pt)
0.5
0.09–0.10
0.924
0.918
clamped ( +0.6 pt)
1.0
0.07–0.08
0.803
0.783
clamped ( +2.0 pt)
1.5
0.02–0.05
0.660
0.645
clamped ( +1.5 pt)
≥2.0
vanished
–
0.37 – 0.54
n/a (single regime)
Table 3: Floor threshold and accuracy vs. task noise (study (v), residual, fair protocol, n=15 seeds/noise level). “Threshold” separates clamped ( γf≈0 ) from peaked ( γf≈4 ); “Winner” has higher test acc. (Appendix D.2 ).
Figure 2: Real-pipeline architecture. An encoder-only upstream model identifies a sweet-spot window among five candidates. That window is then oversampled when training a separate encoder-decoder predictor. The exact upstream score is derived in Appendix F.1 .
Arm
Mean MSE
SD
sweetspot ( μ∗ , ours)
92.4
82.0
uniform (no oversampling)
104.0
82.5
mc_var (BALD-style)
106.5
91.7
random
116.1
85.3
raw_var
139.9
115.6
Table 4: Bacillus, all five windows pooled (13 evaluation points: 5 held-out synthetic + 5 video1 + 3 video2 windows; lower MSE is better).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Synthetic landscape: objective is blind to real signal under H(pA) alone, but not under the residual form.
Family
Objective
obj(info.)
obj(else.)
ratio
γ (i/e)
Rational (main)
H(pA) alone
+0.0047
+0.0050
0.93
7.43/7.87
Rational (main)
H(pA)−H(pβ)
−0.3641
−0.0134
27.11
0.88/6.58
Sigmoid
H(pA) alone
+1.1109
+1.3517
0.82
6.52/6.37
Sigmoid
H(pA)−H(pβ)
−0.2502
−0.2502
1.00
0.16/0.16
Log
H(pA) alone
+0.0042
+0.0046
0.91
7.33/7.74
Log
H(pA)−H(pβ)
−0.4139
−0.0175
23.67
0.86/7.05
Appendix
Table 5: Full functional-form and aggregation-scheme generality table (study ii).
Objective
obj(info.)
obj(else.)
ratio
γ (i/e)
H(pA) alone (reference)
+0.0047
+0.0050
0.93
7.43/7.87
H(pA)−H(pβ) (ours, reference)
−0.3641
−0.0134
27.11
0.88/6.58
H(pA)+1.0(γ/γmax)2
+0.1004
+0.1137
0.88
5.27/5.65
H(pA) , tighter γ clamp ( γmax=2.0 vs. 20.0 )
+0.3830
+0.4543
0.84
2.00/2.00 ‡
H(pA)+5.0∑t(pβ,t−1/T)2
+1.3691
+2.1906
0.62
0.15/5.51
H(pA)−KL(pA∥pβ)
+0.0040
+0.0048
0.83
7.56/7.91
Appendix
Table 6: Full alternative-regularizer table (study iii), including the KL-divergence variant. Same landscape and protocol as Table 5 (30 seeds × 3 μ -inits).
Two-proportion z -test on raw selection rate: z=2.929 , p=0.003 (significant)
Phase 2 (downstream)
Minority recall
Accuracy
Uniform (no emphasis)
0.8466
0.9635
Residual-selected class, fixed 5×
0.8727
0.9687
H(pA) -selected class, fixed 5×
0.8670
0.9675
Appendix
Table 7: Digits ( T=5 ; minority class retained at 8%, approximately 10 training images; 80 seeds and five μ -initializations per objective). Phase 1 reports the true-minority selection rate relative to H(pA) alone. Phase 2 reports minority recall and overall accuracy after separately retraining with a fixed 5× weight on the class selected by each objective. The 2% variant appears in Appendix C.3 .
H(pA) alone
Residual (ours)
floor
lower bdd. (of 6)
mean γf
lower bdd. (of 6)
mean γf
clip effect
0.00
4.00
6.76
0.00
0.01
literal freeze
0.05 – 0.07
5.00
6.77
0.00
0.01
trickle
0.10 – 0.30
5.00
6.79
4.00
≈4.06
trickle
1.00
5.00
6.82
4.00
4.08
trickle
Appendix
Table 8: Study (v), floor sweep under the fair protocol (one μ -init per candidate window, best-of-6 by lowest final objective, 9 seeds per floor, ETRUE=2 , noisy hard-task variant). An expert is lower-bounded when its learning-rate multiplier lr_multe=clip(βe/βˉ,floor,5.0) sits at the lower bound, i.e. lr_multe=floor . Only floor =0 is literal starvation (exact zero gradient). For floor >0 these experts still train, rate-capped at floor.
Noise
static_low
static_high
low → high@75
high → low@75
Winner
0.0
0.957
0.963
0.959
0.963
high ( +0.5 pt)
0.5
0.922
0.916
0.922
0.916
low ( +0.5 pt)
1.0
0.804
0.783
0.804
0.783
low ( +2.1 pt)
1.5
0.644
0.643
0.644
0.643
low ( +0.1 pt)
2.0
0.537
0.537
0.537
0.537
tie
2.5
0.442
0.441
0.442
0.441
low ( +0.1 pt)
Appendix
Table 9: Study (v), floor-adjustment strategy comparison across seven noise levels (residual objective, fair protocol, n=25 per cell). floor low=0.05 , floor high=1.0 . Both schedules switch at step 75 of 150 .
Lower bdd. (of 6)
Ensemble test acc.
H(pA) alone
5.00
0.9628±0.0092
Switch-style aux ( α=0.01 )
5.00
0.9631±0.0094
Residual (ours)
0.00
0.9570±0.0103
Appendix
Table 10: Study (v) extended with a Switch-Transformer-style load-balancing loss (App. D.3 ), 6-candidate fair protocol (25 seeds, best-of-6 by lowest final objective), same protocol as Table 8 .
α
Lower bdd. (of 6)
0.01
5.00
0.1
5.00
1.0
0.00
10.0
0.00
Appendix
Table 11: Load-balancing coefficient sweep, 6-candidate fair protocol (10 seeds each, best-of-6 by lowest final objective), floor =0.05 (never 0 , so counts below are lower-bounded, not literal starvation). Lower-bounding is fixed monotonically by a large enough α .
H(pA) alone
Residual (ours)
floor
lower bdd. (of 6)
mean γf
lower bdd. (of 6)
mean γf
clip effect
0.000
4.00
6.76
0.00
0.01
literal freeze
0.062
5.00
6.77
0.44
0.46
trickle
0.070
5.00
6.77
3.56
3.61
trickle
0.100 – 1.000
5.00
6.78 – 6.82
4.00
≈4.07 – 4.08
trickle
Appendix
Table 12: Study (vi), floor sweep under the fair protocol (one μ -init per candidate window, best-of-6 by lowest final objective, 9 seeds per floor, ETRUE=2 , noisy hard-task variant): lower-bounded-candidate count and mean converged γf , both objectives. Same lower-bounded/literal-starvation distinction as Table 8 (floor =0 only).
H(pA) alone
Residual (ours)
Default (floor =0.05 , clean)
0.960±0.008
0.960±0.010
Hard variant (floor =0 , noisy)
0.756±0.016
0.805±0.017
Floor =1 (clean)
0.960±0.008
0.962±0.007
Appendix
Table 13: Study (vi), fair protocol (one init per window, best-of-6 by lowest final objective, 9 seeds). Same three floor/noise conditions as study (v)’s corresponding ablation (Appendix D.2 ).
gap
mean diff
95% CI
p
dz
0.00
−0.0029
[−0.0056,−0.0003]
0.042
−0.30
0.05
+0.0069
[+0.0035,+0.0103]
0.0003
+0.55
0.10
+0.0144
[+0.0108,+0.0180]
<10−4
+1.08
0.20
+0.0358
[+0.0314,+0.0404]
<10−4
+2.20
0.35
+0.0756
[+0.0704,+0.0810]
<10−4
+3.88
0.50
+0.1414
[+0.1359,+0.1470]
<10−4
+7.04
Appendix
Table 14: Ordering payoff as a function of candidate interchangeability: invest in the chosen expert, then deploy it. Results are paired within run. 25 seeds per cell, floors pooled ( n=50 ). dz is the paired effect size relative to its across-seed variation. Values around 1 or larger are conventionally large.