cs.LGSep 30, 2026

Reliability-Aware Checkpoint Selection for Domain Generalization

Authors: Jinshi Liu, Jiahao Li, Pan Liu, Yanfeng Li, Rui Qian, Zhao Tong, Yue Sun, Tao Tan

Organizations: Shenzhen University · Xiamen University · The Hong Kong University of Science and Technology (Guangzhou) · Macao Polytechnic University · Fudan University · Institute of Information Engineering, Chinese Academy of Sciences

Abstract

Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We identify an empirical selection opportunity within fixed training trajectories: reselecting among checkpoints with near-optimal source accuracy can improve mean target probability quality with small observed changes in mean target accuracy. We study accuracy-constrained reliability selection (AC), which retains checkpoints within a tolerance of the best source-validation accuracy and ranks them by source reliability. Our reference rule aggregates within-set normalized negative log-likelihood (NLL) and class-wise calibration error (CwECE) using D∞D_\infty. AC uses no target data and requires neither additional training nor weight averaging. We evaluate five domain generalization training algorithms on three benchmarks, using PACS to develop the objectives and a 0.5-percentage-point tolerance. In exploratory aggregation comparisons on 360 OfficeHome and TerraIncognita runs, the reference rule reduces mean target soft-bin squared-gap ECE and CwECE by 0.240% and 0.182%, respectively, and NLL by 0.030 relative to Source-Acc. Mean target accuracy changes by +0.213 percentage points. These results identify opportunities for reliability-aware reselection, while the additional benefit of joint over single-objective ranking remains unresolved.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Chaos Is a LADDER: Domain Generalization Beyond Invariance via Reweighting

    Jul 29, 2026Yuhang Jiang, Fengchuan Zhang, Sanguo Zhang +1Multimodal Domain GeneralizationReweighting

  2. ReliableNet: A Chance-Constrained Approach to Trustworthy Classification in Deep Learning

    Aug 10, 2026Ange-Clément Akazan, Ineza Remy Mugenga, Abebe Geletu +2Empirical Risk MinimizationDeep Learning

  3. TopoGeoScore: A Self-Supervised Source-Only Geometric Framework for OOD Checkpoint Selection

    May 9, 2026Farid Hazratian, Ali Zia, Hien Duy NguyenOut-Of-DistributionDomain Invariance