cs.LGOct 8, 2026

Ranking Prior Alignment for Credit Risk Modeling: When Do External Priors Matter?

Authors: Qiye Lu, Jiang Ji, Liang Zhang

Organizations: Ant International China

Abstract

Cold-start credit scoring -- deploying models with scarce labeled data, weak features, or minimal capacity -- is a recurring problem in financial machine learning. When a new lending product launches, labeled default data is scarce, feature pipelines are immature, and models must be deployed with minimal capacity to avoid overfitting. Standard defenses operate on the same limited data; what is needed is a source of external regularization grounded in domain knowledge. We propose Ranking Prior Alignment, a model-agnostic framework that distills external ranking priors (from domain experts, teacher models, or LLMs) into any scoring model via a temperature-scaled KL divergence loss. The framework unifies neural (MIL attention) and tree-based (XGBoost custom objective) architectures through a single formulation: L = L_task + gamma(t) * KL(P_agent || P_model), where gamma(t) follows an exponential decay schedule. The method requires no external model at inference, and its tree-based instantiation tolerates annotation noise up to eta = 0.5. On an industrial dataset of over 1.5M merchants, MIL alignment achieves 7/7 positive evaluation cells at 3K bags (1 ID + 3 OOT + 3 degradation metrics; peak Delta AUC = +0.020 on OOT-1), and XGBoost ablation achieves 9/9 positive metrics at 300 bags. Cross-dataset validation on public Amex shows 5/5 positive folds (avg Delta AUC = +0.041). Four model families (MIL, XGBoost, LightGBM, Logistic Regression) and four teacher architectures show that the framework is both model-agnostic and prior-source-independent. We further observe that alignment gains exhibit an inverse-scaling pattern: benefits grow as data abundance N, model capacity C, and feature quality Q decrease, helping practitioners decide when to invest in prior annotation.

Figures & tables

Explore similar work

CardsList
  1. Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk

    Aug 8, 2026Gregorius Reynaldi Pratama, Kuo-Kun TsengFeature AttributionFaithfulness of Language Model Explanations

  2. Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

    Jan 30, 2026Ruoyu Chen, Shangquan Sun, Xiaoqing Guo +8Shortcut Learning

  3. Privacy-Preserving Credit Risk Prediction with Alternative Data

    Jun 9, 2026Hongzhe Zhang, Jiarong Xu, Jing He +1AI Risk ManagementPrivacy-Preserving ML