cs.LGOct 1, 2026

MECHVAR: Variance-Guided Mechanism Discrimination for Autonomous Machine Learning Experiment Selection

Authors: Yifan Guo

Organizations: Stony Brook University

Abstract

Benchmark gains are often mechanism-ambiguous: reproducing an improvement does not by itself identify why it occurs. We study finite-library mechanism discrimination, where posterior-weighted candidate mechanisms, executable probes, and a limited experimental budget define a sequential experiment-selection problem. MECHVAR selects the next probe by maximizing the posterior-weighted variance of its predicted responses. Under a shared-Gaussian predictive model, this score is exactly proportional to the classical Box--Hill posterior-weighted pairwise-KL criterion, yet it admits O(KE) vectorized rescoring and a transparent additive audit over mechanism pairs. A local expansion further links the score to expected information gain (EIG) when predicted response separations are small. In a 25-block stress audit, MECHVAR outperforms confirmation-first in several moderate misspecification regimes, while its primary comparisons with EIG remain statistically unresolved. In a held-out Digits loop, normalized mechanism-identification AUC is 0.8975 for MECHVAR, 0.7825 for a score-greedy policy, and 0.9092 for EIG. At K = 100, E = 200, median single-thread full-library scoring is 10.36 microseconds for MECHVAR versus 57.69 ms for six-node quadrature EIG in the recorded environment. MECHVAR therefore provides a lightweight, auditable acquisition rule for finite-library experiment selection when a shared predictive scale is a defensible approximation.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer

    Jul 19, 2026Jingjie Ning, Xiaochuan Li, Shanshan Zhong +2Research AutomationAuto Research

  2. Test-Time Scaling via Budgeted Multi-Attribute Verification

    Sep 28, 2026Bo Xue, Ji Cheng, Shen-Huan Lyu +2Token Budget AllocationLarge Language Model Responses

  3. Experimental Experience Modeling for Autonomous Research

    Sep 30, 2026Wenda Wei, Yingchen Zhang, Ruqing Zhang +3Bayesian Experimental DesignPrioritized Experience Replay