Joint Surrogate Learning of Objectives, Constraints, and Sensitivities for Efficient Multi-objective Optimization of Neural Dynamical Systems
Organizations: Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign, Urbana, IL · The Grainer College of Engineering, University of Illinois Urbana-Champaign, Urbana, IL · Department of Neurosurgery, Stanford University, Stanford, CA · Carl R. Woese Institute for Genomic Biology, University of Illinois Urbana-Champaign, Urbana, IL · Mechanical Science and Engineering, University of Illinois Urbana-Champaign, Urbana, IL · National Center for Supercomputing Applications, University of Illinois Urbana-Champaign, Urbana, IL
Abstract
Gaussian process surrogates dominate constrained multi-objective optimization because they are effective in data-scarce regimes, but their cubic scaling in training samples limits their ability to capture shared structure between objectives and constraints as problems grow in dimensionality. We show that deterministic neural network surrogates, equipped with feature tokenization and adaptive output normalization, match or exceed Gaussian process accuracy, while scaling to high-dimensional output spaces and training on all data including infeasible samples. Jointly training a single Feature Tokenizer Transformer to predict objectives, constraint satisfaction, and parameter sensitivities yields a unified gradient that simultaneously improves objective values, steers toward feasibility, and identifies the most influential parameters: a coherent search signal that disjoint per-output models cannot provide. We validate this on biophysical neural optimization problems of increasing complexity. In the hardest regime, with wide, uninformed parameter bounds where random sampling finds zero feasible solutions, descending the surrogate's learned constraint gradient steers the search into the feasible region and recovers near-optimal solutions where standard surrogate optimization and constrained Bayesian optimization find none.
Figures & tables
| Condition | ||
|---|---|---|
| Max blocks | 2 | 2 |
| Max embedding dim | 128 | 96 |
| Max heads | 4 | 3 |
| Min group size | 2 | 4 |
| Pool every blocks | 1 | 1 |
| Constraint BCE | Constraint AUROC | Objective NRMSE | |
|---|---|---|---|
| 0.25 | |||
| 0.5 | |||
| 1 (unweighted) | |||
| 2 | |||
| 4 |
| Category | Parameter | FT-Transformer | ResNet |
|---|---|---|---|
| Architecture | Number of blocks / | 3 | 2 |
| Block dimension / | 128 | 192 | |
| Hidden multiplier / | 2.0 | 2.0 | |
| Number of heads | 4 | – | |
| Embedding dim per head | 32 | – | |
| Pooling | CLS | – |
| Package | Version | Purpose |
|---|---|---|
| Python | 3.12 | Runtime |
| dmosopt | 0.78 (commit f034729 ) | Optimization framework |
| MiV-Simulator | 0.2.1 (commit e440eaa ) | Network simulation |
| NEURON | 8.2.6 | Biophysical neuron simulation |
| TensorFlow | 2.16.2 | Neural network surrogate training |
| PyTorch | 2.5.1 (CPU) | MEGP baseline (via GPyTorch) |
| Parameter | Description | Lower | Upper |
|---|---|---|---|
| pp | Soma area fraction | 0.1 | 0.9 |
| Ltotal | Total length ( m) | 10 | 80 |
| gc | Coupling conductance (mS/cm 2 ) | 0.1 | 50 |
| soma_gmax_Na | Soma Na + conductance (S/cm 2 ) | 0.001 | 0.9 |
| soma_gmax_K | Soma K + conductance (S/cm 2 ) | 0.001 | 1.0 |
| soma_g_pas | Soma leak conductance (S/cm 2 ) | 0 | 0.01 |
| Population | range | range | f-I steps | # params | |
|---|---|---|---|---|---|
| Parvalbumin basket cells (PVBC) | [53, 179] | [5, 21] | 6 | 12 | |
| CCK basket cells (CCKBC) | [240, 320] | [22, 29] | 4 | 12 | |
| Ivy cells (IVY) | [167.5, 367.2] | [50, 70] | 4 | 12 | |
| OLM cells (OLM) | [500.5, 600.1] | [30, 50] | 4 | 12 | |
| Axo-axonic cells (AAC) | [65, 179] | [10, 14] | 5 | 12 | |
| Bistratified cells (BS) | [30.5, 109.1] | [11, 15] | 3 | 12 |
| Narrow range | Wide range | ||||
|---|---|---|---|---|---|
| Parameter | Description | Lower | Upper | Lower | Upper |
| gc | Coupling conductance | 0.1 | 2 | 0 | 2 |
| soma_gmax_Na | Somatic Na + | 0.1 | 0.3 | 0 | 1 |
| soma_gmax_K | Somatic K + | 0.01 | 0.3 | 0 | 1 |
| soma_gmax_KCa | Somatic KCa | 0.0001 | 0.01 | 0 | 1 |
| soma_gmax_CaN | Somatic CaN | 0.00001 | 0.03 | 0 | 1 |
| Population | Slice count |
| Pyramidal (PYR) | 7,825 |
| Parvalbumin basket cells (PVBC) | 142 |
| CCK basket cells (CCKBC) | 104 |
| Ivy cells (IVY) | 217 |
| Neurogliaform cells (NGFC) | 93 |
| OLM cells (OLM) | 43 |
| Experiment | Figures | Trials |
| CA1 single-cell surrogate optimization | 2 , 3 , S2 , S3 | 3 |
| CA1 no-surrogate baseline | 3 , S3 | 3 |
| CA1 optimizer comparison (NSGA-II, AGE-MOEA, SMPSO) | – (nadir estimation) | 3 |
| CA1 sensitivity analysis (surrogate gradient, FAST, DGSM) | 3 , S5 | 5 (HV), 3 (IGD) |
| CA1 initial sampling strategy comparison | 2 , S1 | 3 |
| CA1 gradient-target comparison (BS) | S4 | 3 |
| Figure | Panel / plot | Centre, error, |
|---|---|---|
| 2 B | Sampling strategy comparison (HV) | Mean SEM; dots: individual runs; per strategy (9 populations 3 trials) |
| 2 C | Additive -indicator | Bars mean SEM; dots: individual runs (runs beyond the axis drawn at its edge); per method |
| 2 D | NRMSE / MAPE box plots | Median, IQR box, 1.5 IQR whiskers, dots individual fits within whiskers (fliers not drawn); – per model after outlier filtering of 81 (9 populations 9 checkpoints, one run) |
| 2 E | Inference time | Box plot as in D with individual timings; timed predictions per model |
| 3 A | Convergence curves (HV vs. epoch) | Mean, thin lines individual trials; trials; inset IGD mean SEM with dots |
| 3 B | Mean rank plot | Mean SEM over problems; dots: per-problem ranks |
| gpr | megp | o | c+o | Mean | |
|---|---|---|---|---|---|
| NGFC | |||||
| BS | |||||
| SCA | |||||
| OLM | |||||
| PVBC | |||||
| AAC |
| c+o-FTTransformer | c+o-ResNet | ||||
| Constraint | Violations | AUROC | MCC | AUROC | MCC |
| Firing behaviour | |||||
| Monotonic f-I | 10704 (58.5%) | 0.875 | 0.622 | 0.847 | 0.581 |
| First ISI | 11619 (63.5%) | 0.881 | 0.652 | 0.855 | 0.616 |
| ISI adaptation | 11746 (64.2%) | 0.875 | 0.638 | 0.853 | 0.610 |
| Pre-stimulus spike count | 314 (1.7%) | 0.367 | 0.000 | 0.583 | 0.000 |
| Constraint pairs | Joint surrogate | One model per constraint |
|---|---|---|
| Within firing behaviour | ||
| Within measurement validity | ||
| Across the two groups |
| Model | IVY | PVBC | CCKBC | AAC | IS | SCA | BS | OLM | NGFC |
|---|---|---|---|---|---|---|---|---|---|
| No surrogate | |||||||||
| GPR | |||||||||
| MEGP | |||||||||
| o -ResNet | |||||||||
| c+o -ResNet | |||||||||
| o -FTTransformer |
| Model | IVY | PVBC | CCKBC | AAC | IS | SCA | BS | OLM | NGFC |
|---|---|---|---|---|---|---|---|---|---|
| No surrogate | – | – | – | – | – | – | – | – | – |
| GPR | |||||||||
| MEGP | |||||||||
| o -ResNet | |||||||||
| c+o -ResNet | |||||||||
| o -FTTransformer |
| Model | IVY | PVBC | CCKBC | AAC | IS | SCA | BS | OLM | NGFC |
|---|---|---|---|---|---|---|---|---|---|
| No surrogate | |||||||||
| GPR | |||||||||
| MEGP | |||||||||
| o -ResNet | |||||||||
| c+o -ResNet | |||||||||
| o -FTTransformer |
| Model | IVY | PVBC | CCKBC | AAC | IS | SCA | BS | OLM | NGFC |
|---|---|---|---|---|---|---|---|---|---|
| MEGP (feasible only) | |||||||||
| MEGP (all samples) | |||||||||
| Rank-1 LMC | – |
| Model | IVY | PVBC | CCKBC | AAC | IS | SCA | BS | OLM | NGFC |
|---|---|---|---|---|---|---|---|---|---|
| MEGP (feasible only) | |||||||||
| MEGP (all samples) | |||||||||
| Rank-1 LMC | – |
| Model | IVY | PVBC | CCKBC | AAC | IS | SCA | BS | OLM | NGFC |
|---|---|---|---|---|---|---|---|---|---|
| MEGP (feasible only) | |||||||||
| MEGP (all samples) | |||||||||
| Rank-1 LMC | – |
| Model | Fit (s) | Prediction (s) | Fit slowdown | Prediction slowdown |
|---|---|---|---|---|
| MEGP (feasible only) | – | – | ||
| MEGP (all samples) |
| All test points | Feasible test points | |||||
|---|---|---|---|---|---|---|
| Epoch | Training set | NRMSE | Spearman | NRMSE | Spearman | |
| 5 | MEGP (feasible only) | 372 | [0.32, 0.56] | [0.13, 0.27] | [0.21, 0.50] | [0.34, 0.54] |
| MEGP (all samples) | 1533 | [0.20, 0.67] | [0.30, 0.55] | [1.44, 3.68] | [0.20, 0.44] | |
| 13 | MEGP (feasible only) | 607 | [0.40, 1.15] | [0.23, 0.37] | [0.34, 1.20] | [0.45, 0.58] |
| MEGP (all samples) | 2333 | [1.47, 9.90] | [0.28, 0.52] | [4.30, 82.92] | [-0.02, 0.28] | |
| Narrow range (MN) | Wide range (MN-r) | |||
| Method | HV-AUC | Final HV | HV-AUC | Final HV |
| Constrained BO ( ) | † | † | † | † |
| Constrained BO ( ) | ||||
| Constrained BO ( ) | ||||
| GPR | ||||
| MEGP | ||||