Soft Strategy Selection for Batch-Mode Active Learning
Authors: Rushil Gupta, Romain Lopez
Organizations: Department of Computer Science, Courant Institute School of Mathematics, Computing, and Data Science, New York University · Department of Biology, New York University
Real-world deployment of active learning typically forces practitioners to choose an acquisition strategy before any data is labeled. This is a daunting task: strategy performance varies widely across settings (e.g. datasets, surrogate models) and cannot be assessed without deployment. Existing strategy selection methods explore one strategy from a portfolio at each round and identify the optimal one using bandit feedback or model retraining. Many acquisition rounds are therefore spent exploring strategies rather than collecting the most informative data. Such overhead is a major barrier to AL-driven design of high-throughput experiments, such as genetic perturbation screens and directed evolution, where AL runs consist of only a few rounds with large batch sizes. This regime permits a natural alternative: acquiring data using multiple AL strategies within a single batch. We refer to this as soft strategy selection and introduce FractAL, a method specifically designed for this task. FractAL infers a per-strategy reward using influence-function-based data attribution, which requires no additional retraining, and then computes budget shares for each strategy in the portfolio using online mirror descent. We benchmark FractAL across 7 setups spanning classification, regression, and genetic perturbation effect prediction. The results highlight that strategy selection is a hard problem: every existing method performs worse than random sampling on at least one setup. FractAL, however, matches or outperforms every baseline, including random sampling, on all 7 setups. Its allocations concentrate budget on the strongest strategies in the portfolio while pruning the weakest. FractAL is therefore a reliable choice for real-world deployments, where the optimal strategy is unknown in advance, an important step towards making AL practical for high-throughput experiments and modern scientific discovery.
Figures & tables
Figure 1: Performance gain over random sampling for the six adaptive strategy selection methods, across the seven setups of Section 4 (Appendix G ).
Related work
Batch Sampling
Fractional Budget Allocation
No Retraining / Aux. Models
ALBL ( Hsu and Lin, 2015 )
✗
✗
✓
AutoAL ( Wang et al., 2025 )
✓
✗
✗
SelectAL ( Hacohen and Weinshall, 2023 )
✓
✗
✗
FractAL (this work)
✓
✓
✓
Table 1: Comparison with related work on three axes: support for batch sampling, fractional budget allocation across the strategy portfolio, and the ability to operate without additional model training.
Figure 2: Setup (C1): CIFAR-10. (a) Test accuracy of the three individual strategies. (b–d) Per-round average budget allocation under FractAL, SelectAL, and ALBL. Solid lines denote the mean and shaded bands indicate one standard deviation.
Table 4
Chance
SCRiBLe
SelectAL
FractAL
Setup
Best-Set Overlap
Worst-Set Overlap
Best-Set Overlap
Worst-Set Overlap
Best-Set Overlap
Worst-Set Overlap
Best-Set Overlap
Worst-Set Overlap
CIFAR-10 (C1)
0.33
0.33
0.34 ± 0.03
0.35 ± 0.05
0.46 ± 0.02
0.41 ± 0.03
0.79 ± 0.02
0.89 ± 0.02
CIFAR-10 (C2)
0.25
0.13
0.24 ± 0.01
0.10 ± 0.02
0.29 ± 0.02
0.17 ± 0.02
0.46 ± 0.01
0.61 ± 0.01
KEGG
0.50
0.25
0.51 ± 0.01
0.29 ± 0.03
0.50 ± 0.01
0.24 ± 0.01
0.58 ± 0.01
0.31 ± 0.02
SARCOS
0.50
0.75
0.51 ± 0.01
0.75 ± 0.01
0.53 ± 0.02
0.77 ± 0.01
0.54 ± 0.01
0.79 ± 0.01
Diamonds
0.50
0.50
0.52 ± 0.02
0.52 ± 0.01
0.51 ± 0.01
0.51 ± 0.01
0.52 ± 0.02
0.52 ± 0.02
Table 4: Best-Set Overlap ( ↑ ) and Worst-Set Overlap ( ↑ ) across all seven setups. Chance is the expected overlap under a uniform allocation, ∣set∣/K (set cardinalities in Table 6 ).
Method
(B1): BMDMs
(B2): T cells
Pure strategies (best/worst in hindsight)
Best pure x
0.4377 ± 0.0005 CoreSet
0.3696 ± 0.0002 CoreSet
Worst pure x
0.4165 ± 0.0011 Random
0.3403 ± 0.0045 Random
Non-adaptive controls
Random
0.4165 ± 0.0011
0.3403 ± 0.0045
Equal Split
0.4279 ± 0.0006
0.3442 ± 0.0065
Table 5: AUPRC ( ↑ ) for perturbation prediction tasks.
Figure 3: Setup (B2): T cell perturbation prediction. (a) Test AUPRC of the four pure strategies. (b–d) Per-round average budget allocation under FractAL, SelectAL and ALBL. Solid lines denote the mean and shaded bands indicate one standard deviation.
Figure 4: Runtime for an AL run on (C2).
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Bt←Sample(Akt,B,Dpool∖Dtraint−1)
Appendix
Algorithm 2 ALBL: Active Learning by Learning
kt=argmaxk(μ^k+cnklogt)
Appendix
Algorithm 3 UCB: Upper Confidence Bound
Dtraint←Dtraint−1∪{(x,y):x∈Bt}
Appendix
Algorithm 4 SCRiBLe
Dataset
∣Best Set∣
∣Worst Set∣
CIFAR 10 (C1)
1
1
CIFAR 10 (C2)
2
1
KEGG
2
1
SARCOS
2
3
Diamonds
2
2
BMDMs (B1)
1
2
Appendix
Table 6: Cardinality of the Best and Worst Sets for the allocation quality assessment for all the datasets
Dataset
Surrogate Model
∣Dinit∣
∣Dval∣
∣Dtest∣
∣Dpool∣
B
CIFAR 10
ResNet-18
500
1024
10000
48976
500
KEGG
3 layer MLP
500
1024
12921
50663
500
SARCOS
3 layer MLP
500
1024
8896
34564
500
Diamonds
3 layer MLP
500
1024
10788
42128
500
BMDMs
1 layer MLP
35
40
103
371
35
CD4+ T cells
1 layer MLP
35
40
141
528
35
Appendix
Table 7: Active learning configurations for each setup: surrogate model, sizes of the initial labeled, validation, test and pool sets, and batch size B . All strategy selection methods and pure strategies share the same configuration within a setup.
Method
Mean Gain ( ↑ )
Worst Gain ( ↑ )
Mean Rank ( ↓ )
Median Rank ( ↓ )
Worst Rank ( ↓ )
Equal Split
0.230 ± 0.214
-0.888
4.29 ± 0.77
4
7
ALBL
0.150 ± 0.227
-1.055
5.43 ± 0.27
6
6
UCB
0.228 ± 0.214
-1.055
4.14 ± 0.55
4
6
SCRiBLe
0.266 ± 0.208
-0.944
4.14 ± 0.47
4
6
SelectAL
0.370 ± 0.178
-0.611
3.14 ± 0.43
3
5
AutoAL
0.027 ± 0.251
-1.055
5.71 ± 0.77
7
7
Appendix
Table 8: Relative gain and rank among the six adaptive methods and the equal split control, averaged across the seven setups. Mean Rank and Worst Rank report the average and maximum rank among the 7 methods (rank 1 = best, computed per setup then averaged/maximized across setups). Median Rank corresponds to the median value of rank across the setups. Mean Gain and Rank values are mean ± SE across setups.
Method
(C1): CIFAR-10 ( K=3 )
(C2): CIFAR-10 ( K=8 )
Individual Strategies
Random
69.4 ± 0.1
69.4 ± 0.1
CoreSet
58.7 ± 0.2
58.7 ± 0.2
ProbCover
71.2 ± 0.1
71.2 ± 0.1
BADGE
69.6 ± 0.2
69.6 ± 0.2
LeastConf
69.6 ± 0.1
69.6 ± 0.1
Appendix
Table 9: Accuracy ( ↑ ) on CIFAR-10 at the end of T=10 rounds.
Method
KEGG
SARCOS
Diamonds
Individual Strategies
Random
0.79 ± 0.02
0.79 ± 0.01
0.11 ± 0.01
CoreSet
0.53 ± 0.01
0.73 ± 0.01
0.08 ± 0.01
ProbCover
0.49 ± 0.00
0.72 ± 0.01
0.12 ± 0.01
TypiClust
0.97 ± 0.01
0.77 ± 0.01
0.10 ± 0.01
Strategy Selection Methods
Appendix
Table 10: Test MSE ( ×10−1 , ↓ ) on tabular regression datasets at the end of T=10 rounds.
Method
B1: BMDM
B2: CD4+ T Cell
Individual Strategies
Random
0.4165 ± 0.0011
0.3403 ± 0.0045
CoreSet
0.4377 ± 0.0005
0.3696 ± 0.0002
ProbCover
0.4305 ± 0.0004
0.3441 ± 0.0003
TypiClust
0.4301 ± 0.0006
0.3413 ± 0.0061
Strategy Selection Methods
Appendix
Table 11: AUPRC ( ↑ ) for high-throughput perturbation datasets after T=6 rounds.
Dataset
Metric
α=0.1
α=0.3
α=0.5
α=0.7
α=1.0
CIFAR-10 (C1)
Acc ( ↑ )
70.0 ± 0.2
70.1 ± 0.2
70.2 ± 0.2
70.0 ± 0.2
70.2 ± 0.1
CIFAR-10 (C2)
Acc ( ↑ )
70.0 ± 0.2
69.9 ± 0.1
69.8 ± 0.2
70.1 ± 0.2
69.9 ± 0.1
KEGG
MSE ( ↓ )
0.60 ± 0.01
0.60 ± 0.01
0.58 ± 0.01
0.58 ± 0.01
0.59 ± 0.01
SARCOS
MSE ( ↓ )
0.73 ± 0.01
0.76 ± 0.01
0.74 ± 0.01
0.74 ± 0.01
0.74 ± 0.01
Diamonds
MSE ( ↓ )
0.09 ± 0.01
0.09 ± 0.01
0.08 ± 0.01
0.09 ± 0.01
0.10 ± 0.01
BMDMs
AUPRC ( ↑ )
0.4348 ± 0.0008
0.4356 ± 0.0006
0.4363 ± 0.0006
0.4360 ± 0.0008
0.4365 ± 0.0008
Appendix
Table 12: Terminal Performance of FractAL for different values of the EWMA smoothing parameter α on all seven setups. MSE numbers represent (MSE ×10−1 )
Dataset
Metric
Mirror Desc.
Additive (clip + renorm.)
CIFAR-10 (C1)
Acc ( ↑ )
70.2 ± 0.2
69.8 ± 0.2
CIFAR-10 (C2)
Acc ( ↑ )
69.8 ± 0.2
69.8 ± 0.1
KEGG
MSE ( ×10−1 , ↓ )
0.58 ± 0.01
0.61 ± 0.01
SARCOS
MSE ( ×10−1 , ↓ )
0.74 ± 0.01
0.76 ± 0.01
Diamonds
MSE ( ×10−1 , ↓ )
0.08 ± 0.01
0.09 ± 0.01
BMDMs
AUPRC ( ↑ )
0.4363 ± 0.0006
0.4299 ± 0.0005
Appendix
Table 13: Performance comparison of mirror descent vs additive update with clipping and renormalization across all setups.
Dataset
Metric
Influence
Retraining
CIFAR-10 (C1)
Acc ( ↑ )
70.2 ± 0.2
70.5 ± 0.1
CIFAR-10 (C2)
Acc ( ↑ )
69.8 ± 0.2
70.2 ± 0.2
KEGG
MSE ( ×10−1 , ↓ )
0.58 ± 0.01
0.59 ± 0.01
SARCOS
MSE ( ×10−1 , ↓ )
0.74 ± 0.01
0.74 ± 0.01
Diamonds
MSE ( ×10−1 , ↓ )
0.08 ± 0.01
0.08 ± 0.01
BMDMs
AUPRC ( ↑ )
0.4363 ± 0.0006
0.4383 ± 0.0006
Appendix
Table 14: Performance comparison with influence vs retraining based reward across all setups.
Round
Spearman Rank Correlation
Best Agreement (%)
Worst Agreement (%)
2
0.70 ± 0.08
70.0
80.0
3
0.80 ± 0.06
80.0
83.3
4
0.78 ± 0.04
70.0
86.6
5
0.73 ± 0.08
80.0
76.6
6
0.73 ± 0.07
76.6
76.6
Appendix
Table 15: Correspondence between the retraining based reward and corresponding influence estimates across different active learning rounds for the CIFAR-10 (C1) setup.
Dataset
Metric
FractAL
FractAL (Hard)
CIFAR-10 (C1)
Acc ( ↑ )
70.2 ± 0.2
69.3 ± 0.1
CIFAR-10 (C2)
Acc ( ↑ )
69.8 ± 0.2
70.0 ± 0.1
KEGG
MSE ( ×10−1 , ↓ )
0.58 ± 0.01
0.60 ± 0.01
SARCOS
MSE ( ×10−1 , ↓ )
0.74 ± 0.01
0.74 ± 0.01
Diamonds
MSE ( ×10−1 , ↓ )
0.08 ± 0.01
0.09 ± 0.01
BMDMs
AUPRC ( ↑ )
0.4363 ± 0.0006
0.4309 ± 0.0010
Appendix
Table 16: Performance comparison of hard versus soft allocation across all setups.
Method
Acquisition
Allocation Update
Total
ALBL
1.2
7.3 ×10−5
1.2
UCB
1.5
7.2 ×10−5
1.5
SCRiBLe
8.8
4.6 ×10−4
8.8
SelectAL
3.3
19.8
23.1
AutoAL
20.7
28.5
49.2
FractAL
8.8
0.5
9.3
Appendix
Table 17: Runtime (in minutes) of different components of a single AL run for different methods. We use the CIFAR10 (C2, K=8 ) setting for these measurements.
Figure 5: Setup (C1): CIFAR-10 with K=3 . (a) Test performance of pure strategies. (b–f) Average per-round budget allocation to Random, CoreSet, and ProbCover under each method. FractAL drives CoreSet to zero and concentrates on ProbCover. SelectAL allocates roughly equal shares to Random and CoreSet, while bandit baselines (d–f) oscillate between strategies without converging. Lines show the mean and shaded bands show ± 1 standard deviation. For test accuracy, standard deviation is computed across seeds after averaging over splits; for allocations, across splits after averaging over seeds.
Figure 6: Setup (C2): CIFAR-10 with K=8 . (a) Test performance of the 8 pure strategies. (b–f) Average per-round budget allocation to all 8 strategies under each method. FractAL increases the share of ProbCover and decreases CoreSet over rounds. Lines show the mean and shaded bands show ± 1 standard deviation. For test accuracy, standard deviation is computed across seeds after averaging over splits; for allocations, across splits after averaging over seeds.
Figure 7: KEGG metabolic reaction network experiment. (a) Test performance of the four pure strategies. (b–f) Average per-round budget allocation under each method. FractAL concentrates the budget on CoreSet (the second strongest strategy) while suppressing Random and TypiClust. SelectAL also upweights CoreSet but assigns ProbCover a lower share than Random consistently. The bandit baselines (d–f) oscillate between strategies without converging. Lines show the mean and shaded bands show ± 1 standard deviation. For test accuracy, standard deviation is computed across seeds after averaging over splits; for allocations, across splits after averaging over seeds.
Figure 8: SARCOS experiment. (a) Test performance of the four pure strategies. (b–f) Average per-round budget allocation under each method. FractAL and SelectAL upweight CoreSet (one of the strongest strategies) and suppress Random and ProbCover. SelectAL ranks TypiClust at the lowest allocation, despite it not being the weakest strategy; FractAL avoids this mistake. The bandit baselines (d–f) oscillate without converging. Lines show the mean and shaded bands show ± 1 standard deviation. For test accuracy, standard deviation is computed across seeds after averaging over splits; for allocations, across splits after averaging over seeds.
Figure 9: Diamonds experiment. (a) Test performance of the four pure strategies, which cluster within a narrow MSE range. (b–f) Average per-round budget allocation under each method. FractAL upweights CoreSet (the strongest individual strategy) and reduces the others, with the allocations aligned with the empirical ranking by the final rounds. The remaining methods (c–f) maintain roughly uniform allocations. Lines show the mean and shaded bands show ± 1 standard deviation. For test accuracy, standard deviation is computed across seeds after averaging over splits; for allocations, across splits after averaging over seeds.
Figure 10: Setup (B1): BMDM perturbation prediction. (a) Test AUPRC of the four pure strategies. (b–f) Average per-round budget allocation under each method. FractAL and SelectAL both upweight CoreSet (the strongest strategy) over rounds, but FractAL drives the other three strategies close to zero, whereas SelectAL retains substantial allocation to Random. The bandit baselines (d–f) oscillate between strategies without converging. Lines show the mean and shaded bands show ± 1 standard deviation. For test accuracy, standard deviation is computed across seeds after averaging over splits; for allocations, across splits after averaging over seeds.
Figure 11: Setup (B2): CD4+ T cell perturbation prediction. (a) Test AUPRC of the four pure strategies. (b–f) Average per-round budget allocation under each method. FractAL consistently upweights CoreSet (the strongest strategy) and reduces the other three to similar low levels, consistent with their similar pure-strategy performance. SelectAL temporarily reduces CoreSet’s allocation in rounds 3–5, and keeps ProbCover proportions high. The bandit baselines (d–f) oscillate between strategies without converging. Lines show the mean and shaded bands show ± 1 standard deviation. For test accuracy, standard deviation is computed across seeds after averaging over splits; for allocations, across splits after averaging over seeds.
Department of Electrical and Computer Engineering, University of Minnesota, Twin Cities, USA · Jiacong Li, Tianpei Xie, Cecile Levasseur, and Wojciech Kowalinski are with Amazon.