Kernel Autoresearch for Open-Ended Model Discovery
Organizations: CUHK-Shenzhen · University of British Columbia
Abstract
Kernels encode the inductive bias of a wide range of machine learning models, yet automated kernel design faces a fundamental dilemma. A fixed grammar of base kernels and operators guarantees validity but limits the search to structures expressible by those building blocks. Conversely, unrestricted programs remove this limitation but no longer guarantee validity. In our stress tests, 22-58% of LLM-generated kernels that pass numerical checks on random inputs fail when evaluated at different scales or dimensions. We propose Kernel Autoresearch (Kernaut), which treats kernel design as open-ended model discovery. Coding agents write kernels as programs, while construction contracts ensure that every accepted kernel is valid. A quality-diversity archive retains high-performing kernels with distinct behaviors, and novelty screening steers agents toward functionally new candidates. Our experiments demonstrate that the discovered kernels encode reusable inductive biases that generalize to unseen tasks. On held-out black-box optimization families, a discovered kernel outperforms a meta-learned deep kernel trained on the same episodes. Furthermore, kernels discovered from ten enzyme-kinetic rate laws achieve lower error than tuned ARD and deep kernel baselines on five unseen mechanisms. The discovered kernels are also interpretable programs that human researchers can refine: a human-refined version of one further reduces the held-out predictive error by 5.7% and optimization regret by 7.8%.
Figures & tables
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
| Weighting | Weights | DWF rank | Elites kept |
|---|---|---|---|
| Original | 1 | 9/9 | |
| AUC-heavy | 1 | 9/9 | |
| Equal | 1 | 9/9 | |
| CRPS-only | 1 | 9/9 |
| T | Label | What is established |
|---|---|---|
| 0 | Executable | Sandboxed execution succeeds and the output matches the interface |
| 1 | Empirical | Randomized Gram-matrix tests pass on sampled input sets |
| 2 | Contract certified | Checked interface conformance and trusted PSD-preserving assembly under the stated component assumptions |
| Relation | What must match |
|---|---|
| Determinism | The same input, replayed twice |
| Permutation | The Gram of a permuted batch equals the permuted Gram |
| Subset | The Gram of a random subset equals the corresponding submatrix |
| Extension | Appending then truncating points reproduces the original block |
| Duplicate | A duplicated point reproduces its own row |
| Field | Purpose |
|---|---|
| mathematical_form | Kernel expression and construction |
| psd_argument | Proof sketch for positive semidefiniteness |
| novelty_claim | What is new relative to known kernels |
| closest_known_kernel | Nearest baseline and how this differs |
| expected_bo_behavior | Predicted BBO performance characteristics |
| construction_niche | Declared construction category |
| Family | Split(s) | |
| Branin | 2 | train, val |
| Ackley | 2 | train, val |
| Cosine | 8 | train, val |
| Hartmann-6 | 6 | train, val |
| Levy | 6 | train, val |
| Bukin | 2 | test |
| CRPS | AUC | |||
| Candidate | val | test | val | test |
| Linear | 0.783 | 0.791 | 0.913 | 0.872 |
| Periodic ( , ) | 0.549 | 0.618 | 0.816 | 0.678 |
| Rational quadratic (RQ; , ) | 0.559 | 0.618 | 0.785 | 0.706 |
| Matérn-5/2 ( ) | 0.551 | 0.612 | 0.797 | 0.693 |
| RBF ( ) | 0.565 | 0.641 | 0.823 | 0.756 |
| Search | Test CRPS | Mean AUC | Failed calls |
|---|---|---|---|
| Full archive | 0.681 | 28 | |
| Independent proposals | 0.695 | 15 | |
| Closure grammar | 0.706 | 22 | |
| CKS greedy grammar, depth | 0.684 | – |
| Validation ( ) | Test ( ) | |||||
|---|---|---|---|---|---|---|
| Competitor | wins | med. CRPS | wins | med. CRPS | ||
| vs. fixed baselines | ||||||
| Periodic ( , ) | 34/50 | * | 45/60 | *** | ||
| Matérn-5/2 ( ) | 31/50 | 39/60 | * | |||
| RQ ( , ) | 39/50 | *** | 45/60 | *** | ||
| vs. discovered kernels | ||||||
| Construction | Test CRPS (repeat mean SD) | Selected val. | Selected test | Selected AUC |
|---|---|---|---|---|
| Identity Matérn | 0.5546 | 0.6189 | 0.6824 | |
| Warp only | 0.5590 | 0.6223 | 0.6896 | |
| Fold only | 0.5519 | 0.6147 | 0.6901 | |
| Additive branches ‡ | 0.5020 | 0.5668 | 0.6190 | |
| Joint branches | 0.5404 | 0.6011 | 0.6712 | |
| Rational quadratic | 0.5489 | 0.6090 | 0.6896 |
| Construction | Test CRPS | Regret AUC |
|---|---|---|
| Joint | 0.601146 | 0.671159 |
| Joint, smooth | 0.601162 | 0.671159 |
| Additive ‡ | 0.566832 | 0.618955 |
| Additive, smooth ‡ | 0.566879 | 0.618955 |
| Per discovery run | BBO | Time-series |
|---|---|---|
| Rounds | 1–10 | 8–10 |
| Tool calls | 0–11 | 7–10 |
| Input tokens | 4k–190k | 69k–184k |
| Output tokens | 0.3k–10k | 3k–9.6k |
| Method | val | SF 6 |
|---|---|---|
| Linear | 0.826 | 0.959 |
| Periodic ( annual) | 0.241 | 1.799 |
| RQ ( ) | 0.095 | 0.067 |
| Matérn-5/2 | 0.091 | 0.050 |
| Matérn-3/2 | 0.099 | 0.075 |
| RBF |
| Kernel | All CRPS | Forward CRPS | Reversed CRPS | Forward coverage | Forward width |
|---|---|---|---|---|---|
| Residual period-bank (selected) | 0.0456 | 0.0488 | 0.0391 | 0.918 | 0.334 |
| Harmonic–trend root | 0.0194 | 0.0174 | 0.0235 | 0.925 | 0.121 |
| RBF ( ) | 0.0264 | 0.0284 | 0.0223 | 0.880 | 0.153 |
| Uniform ARD-RBF ( ) | 0.0184 | 0.0178 | 0.0195 | 0.916 | 0.117 |
| Method | SF 6 | CFC-12 | CFC-11 | Mean | Fwd. |
|---|---|---|---|---|---|
| Fixed (RBF) | |||||
| Best fixed ∗ | |||||
| GSM | |||||
| Input warping | |||||
| RFF | |||||
| CAKE |
| Comparator | SF 6 | CFC-12 | CFC-11 | Pooled |
|---|---|---|---|---|
| Fixed (RBF, ) | 7/30 ( ) | 26/30 ( ) | 25/30 ( ) | 58/90 ( ) |
| Input warping | 26/30 ( ) | 12/30 ( ) | 15/30 ( ) | 53/90 ( ) |
| Random Fourier features | 14/30 ( ) | 21/30 ( ) | 27/30 ( ) | 62/90 ( ) |
| Grid spectral mixture | 30/30 ( ) | 30/30 ( ) | 29/30 ( ) | 89/90 ( ) |
| CAKE | 2/30 ( ) | 16/30 ( ) | 16/30 ( ) | 34/90 ( ) |
| Best fixed per record ∗ | 7/30 ( ) | 17/30 ( ) | 14/30 ( ) | – |
| Reference (test CRPS) | Ensemble E1 ( ) | Frontier F2 ( ) |
|---|---|---|
| Optimized ARD ( ) | , | , 73/75, |
| Coordinate-aware DKL ( ) | , | , 71/75, |
| Domain library ( ) | , | , 75/75, |
| Monotone warping ( ) | , | , 71/75, |
| Isotropic Matérn ( ) | , | , 74/75, |
| Shared-coordinate DKL ( ) | , | , 74/75, |
| Discovery | b2 c45 | b3 c45 | b4 c45 | b6 c0 | b6 c45 | b6 c120 | b6 c180 |
|---|---|---|---|---|---|---|---|
| Frontier F1 | 10, .002 | 9, .004 | 9, .027 | 10, .002 | 10, .002 | 10, .002 | 6, 1.0 |
| Frontier F2 | 8, .014 | 10, .002 | 7, .027 | 10, .002 | 8, .020 | 9, .014 | 6, .92 |
| Ensemble c2 | 9, .004 | 10, .002 | 8, .037 | 8, .020 | 7, .049 | 8, .014 | 5, .62 |
| Ensemble c1 | 10, .002 | 6, .28 | 5, .77 | 8, .027 | 7, .23 | 7, .56 | 2, .020 |
| Ensemble c3 | 10, .002 | 9, .010 | 7, .19 | 5, 1.0 | 5, .77 | 5, .77 | 2, .037 |
| Kernel | Report | What tuning settled |
|---|---|---|
| DWF (BBO) | The objective varies smoothly, as under a Matérn-5/2 prior, along a slightly rescaled copy of each coordinate. Values at points mirrored about the middle of a coordinate, and , are encouraged but not forced to agree. The prior has a kink at the mirror line . | Tuned geometry: the fold frequency settled at , so the fold is a single tent peaking at that reflects each coordinate once, and the warp stays within of the identity (Appendix C.2 ). The registration had anticipated multi-scale periodicity and boundary sensitivity, which the tuned kernel does not use. The kink is Proposition 4 . |
| Residual period-bank + poly (forecasting) | The record follows a smooth trend, as under a Matérn-5/2 prior with lengthscale of the normalized span, plus a constant offset, plus a second smooth component built from 36 sinusoids. Every sinusoid has a period between 7 and 47 times the length of the record, so within one record this component acts as a record-wide smooth trend rather than a cycle. | Tuned behavior: the sinusoid frequencies settled at – cycles per normalized span, so the bank models slow curvature across the whole record. The registration described an annual cycle, but it took the annual period (about of the span) as a frequency. An annual cycle would need 25–47 cycles per span. The “poly” feature is a constant. |
| ChemBench, ensemble E1 (enzyme kinetics) | The rate depends smoothly on a monotone transform of log substrate concentration that is steepest at high concentration, the same transform of the second substrate, their product, and log enzyme loading. Inhibitor and product concentrations matter only in proportion to substrate saturation. | Tuned relevance: as in ARD, tuning decides which inputs matter. Enzyme loading and the two substrates matter most, while inhibitor and product matter little (Figure 28 ), and the saturation midpoint lies above the input range (Figure 20 ). The program includes Arrhenius and pH-window features, but tuning set their weights so that the correlation is at every base point. These probes show little sensitivity to temperature and pH, without establishing exact invariance. |
| Kernel | Report | What tuning settled |
|---|---|---|
| ChemBench, frontier F2 (enzyme kinetics) | An explicit 248-feature map treats the rate as a linear combination of ten substrate-response shapes modeled on the training mechanisms, each at three substrate scales, their log-substrate sensitivities, and a small linear backbone. The response features are multiplied by enzyme, temperature, and pH factors, so posterior coefficients can select both mechanism shape and environmental modulation. | Tuned relevance: enzyme loading and substrate dominate, product, inhibitor, and temperature matter less, and pH is nearly ignored (Figure 28 ). The frozen kernel therefore retains strong enzyme and substrate effects, some product and inhibitor response, and almost no pH dependence. |
| GlucoseBench, frontier F1 | An ARD-Matérn-5/2 residual is augmented by 36 intervention features. At three response scales, gamma envelopes encode causal meal and bolus transients, long-period wave packets encode response age, and cumulative channels retain dose exposure after the transients fade. Signed and complementary channels allow meal and bolus responses to oppose or covary. | Tuned behavior: signed channels dominate because their mixing parameter is , while cumulative memory remains active with gain . The base kernel’s time lengthscale is normalized spans and its four intervention lengthscales are . On public training inputs, endpoint-sweep correlations are for time, for meal size, for meal start, and for bolus size and start. |