Driven by recent advances in artificial intelligence, a growing literature has demonstrated the potential of using large language models (LLMs) as scalable surrogates to generate human-like responses. Two common approaches to improve the performance of LLMs include: fine-tuning, which aligns the LLM more closely with human responses, and rectification, which corrects biases in LLM outputs. In this paper, we develop a two-stage framework that combines fine-tuning and rectification, and optimally allocates limited labeled samples across the two stages. A key insight is that the conventional fine-tuning objective of minimizing mean squared prediction error is generally not aligned with the downstream rectification stage. For mean estimation, we propose to minimize the variance of the prediction errors; for general M-estimation, we propose to minimize a scalarized variance metric as the fine-tuning objective. Building on this insight, we leverage the scaling law of fine-tuning to optimally allocate the limited labeled human data between the fine-tuning and rectification stages. Our empirical analysis validates the fine-tuning scaling law and confirms that our proposed optimal allocation rule reliably identifies the optimal sample allocation. We demonstrate substantial efficiency gains in estimation and inference performance relative to fine-tuning or rectification alone, or to employing the conventional mean squared error objective within the fine-tuning then rectification framework. Such efficiency gains translate to significant cost savings for making reliable decisions.
Figures & tables
Review
This is a walk backward after the impressive 2012. Almost impenetrably black, the flavors converge around espresso and bitter chocolate, yet the tannins have a green edge. The wine simply feels flat in the mouth with no life to it.
Price
$28
Rating ( Y )
85
Review
This is one of the great Rieslings from the Wachau, a wonderful panoply of ripe, tropical fruit, pierced with flint, spice and minerality. It is rich and opulent, while never losing sight of the core tautness of a fine Riesling.
Price
$70
Rating ( Y )
95
Table 1 : Examples from the Wine Reviews dataset
Figure 1 : Empirical validation of the variance-based scaling law. The estimated parameters are a=11.06 , α=0.318 , and b=2.17 (adjusted R2=0.999 ); points are medians over ten training replicates per size with 95% error bars ( 1.96 standard errors over the replicates)
Figure 2 : Mean estimated 95% CI width of the FT+PPI estimator across sample allocations, for the variance-loss and MSE-loss fine-tuning objectives.
Figure 3 : Relative variance reduction of FT+PPI versus the Sample Mean , computed from the fitted scaling law.
n
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE ( ×10−2% )
1000
Sample Mean
10.062(0.432)
8.018(0.352)
9.060(0.397)
FT-Only
32.462(0.469)
31.308(0.496)
35.380(0.561)
PPI-Only
7.092(0.306)
5.621(0.250)
6.352(0.283)
FT+PPI (MSE Loss)
9.767(0.432)
7.653(0.351)
8.649(0.397)
FT+PPI (Var Loss)
7.155(0.261)
5.819(0.241)
6.575(0.272)
3000
Sample Mean
5.783(0.226)
4.600(0.203)
5.198(0.229)
Table 2 : Estimation accuracy across estimators under different labeled sample sizes n
n
FT+PPI (Var Loss) vs. Sample Mean
FT+PPI (MSE Loss) vs. Sample Mean
FT+PPI: Var vs. MSE Loss
1000
47.84(4.43)
7.83(14.47)
42.27(8.87)
3000
53.02(3.24)
31.22(4.54)
31.46(5.87)
5000
53.91(2.15)
37.05(3.14)
26.65(4.35)
10000
51.90(1.51)
38.36(1.92)
21.91(2.89)
Table 3 : Variance reductions (%) across estimators
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE (%)
OLS
2.436(0.100)
1.956(0.084)
0.997(0.043)
PPI-Only
1.679(0.066)
1.341(0.058)
0.684(0.030)
FT+PPI (MSE Loss)
1.420(0.055)
1.155(0.048)
0.589(0.024)
FT+PPI (Var Loss)
1.404(0.051)
1.145(0.047)
0.584(0.024)
Table 4 : Estimation accuracy for the price-rating regression coefficient at n=10,000
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE (%)
OLS
2.976(0.148)
2.394(0.125)
3.303(0.173)
FT+PPI (MSE Loss)
3.839(0.185)
3.063(0.164)
4.227(0.226)
FT+PPI (Var Loss)
3.259(0.165)
2.599(0.139)
3.586(0.192)
Table 5 : Controlled synthetic study: estimation accuracy for the census slope at n=1,200
Figure 4 : Bootstrap robustness of the scaling-law fit based on 5,000 -sample development pools. Light curves correspond to individual bootstrap replicates, the solid line denotes the median fit, and the shaded region indicates the central 95% robustness interval.
Quantity
Median
95% Robustness Interval
Scaling parameter a
14.265
[10.057,63.832]
Scaling parameter α
0.391
[0.251,0.531]
Variance floor b
2.338
[0.990,2.434]
Optimal allocation s∗/n
0.104
[0.094,0.107]
Table 6 : Robustness of the scaling-law fit
Figure 5 : Validation residual variance under variance-based and MSE-based fine-tuning, with the constant-predictor baseline ( Var(Y) ) for reference; current recipe (frozen text-embedding-3-large + MLP head).
n
RMSE (Cross-entropy) ( ×10−2 )
RMSE (Var Loss) ( ×10−2 )
1,000
10.038
7.873
3,000
4.303
2.567
5,000
3.037
2.816
10,000
2.471
2.514
Table 7 : Cross-entropy training versus the variance loss
Estimator
Residual variance
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE ( ×10−2% )
FT+PPI (MSE Loss)
3.626
2.187
1.743
1.969
FT+PPI (Var Loss)
3.205
1.951
1.580
1.786
Table 8 : LoRA fine-tuning: residual variance and PPI estimation accuracy
Design
Training Objective
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE ( ×10−2% )
Sample Mean
—
3.374(0.121)
2.766(0.112)
3.126(0.126)
PPI-Only
—
2.473(0.104)
1.991(0.085)
2.250(0.096)
FT+PPI split
MSE
2.381(0.094)
1.892(0.084)
2.138(0.094)
FT+PPI split
Variance
2.177(0.082)
1.752(0.075)
1.980(0.084)
Cross-fitting K=3
MSE
1.909(0.083)
1.509(0.068)
1.706(0.076)
Variance
1.734(0.081)
1.372(0.061)
1.551(0.069)
Table 9 : Cross-fitting on the main corpus: training objective at each fold count
n
Training Objective
PPI Variance
PPI++ Variance
Mean λ
1,000
MSE
9.38×10−3
7.58×10−3
0.61
Variance
5.17×10−3
5.06×10−3
0.97
10,000
MSE
6.11×10−4
5.42×10−4
0.82
Variance
4.79×10−4
4.65×10−4
0.95
Table 10 : Estimator variance under plain PPI and PPI++
Zero-Shot Model
Corr. with Human Ratings
Variance reduction w.r.t. Sample Mean
gpt-5.4-nano
0.741
52% – 54%
gpt-5.6-terra
0.860
72% – 74%
gpt-5.6-sol
0.863
73% – 74%
gpt-5.6-luna
0.885
77% – 78%
Table 11 : Zero-shot model quality and estimator efficiency under affine calibration
n
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE (%)
1,000
Sample Mean
3.286(0.328)
2.905(0.352)
0.694(0.084)
PPI-Only
2.293(0.341)
1.886(0.299)
0.451(0.071)
FT+PPI (MSE Loss)
3.728(0.445)
3.079(0.482)
0.736(0.115)
FT+PPI (Var Loss)
2.539(0.463)
2.041(0.347)
0.488(0.083)
2,000
Sample Mean
2.504(0.462)
1.923(0.368)
0.459(0.088)
PPI-Only
1.534(0.135)
1.384(0.152)
0.331(0.036)
Table 12 : Customer-review corpus: estimation accuracy across estimators
n
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE (%)
250
Sample Mean
8.603(1.166)
6.997(1.149)
1.881(0.309)
PPI-Only (prior wave)
5.251(0.591)
4.402(0.657)
1.183(0.177)
FT+PPI (MSE Loss)
8.452(1.055)
7.138(1.039)
1.919(0.279)
FT+PPI (Var Loss)
8.798(1.250)
6.841(1.269)
1.839(0.341)
500
Sample Mean
6.245(0.728)
5.309(0.755)
1.427(0.203)
PPI-Only (prior wave)
4.732(0.669)
3.878(0.622)
1.043(0.167)
Table 13 : Digital-twin survey panel: estimation accuracy across estimators
Component
Role
Value
Target-population size
census size
4,000
Cluster location zhi
center of the high-influence cluster
3.5
Cluster mass
share of the census in the cluster
15%
Cluster signal weight κ
outcome weight of the high-influence component
3.0
Outcome noise σ
homoskedastic noise scale
0.8
Nuisance dimensions
uninformative embedding coordinates
5
Table 14 : Generating parameters of the synthetic design
Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions. We study the design problem of allocating a fixed budget of human respondents across estimation tasks when cheap LLM predictions are available for every task. Our framework combines three components. First, building on Prediction-Powered Inference, we characterize a question-specific rectification difficulty that governs how quickly the estimator's variance decreases with human sample size. Second, we derive a closed-form optimal allocation rule that directs more human labels to tasks where the LLM is least reliable. Third, since rectification difficulty depends on unobserved human responses for new surveys, we propose a meta-learning approach, trained on historical data, that predicts it for entirely new tasks without pilot data. The framework extends to general M-estimation, covering regression coefficients and multinomial logit partworths for conjoint analysis. We validate the framework on two datasets spanning different domains, question types, and LLMs, showing that our approach captures 61-79% of the theoretically attainable efficiency gains, achieving 11.4% and 10.5% MSE reductions without requiring any pilot human data for the target survey.
In Large Language Model (LLM) fine-tuning, parameter and data selection are common strategies for reducing fine-tuning cost, yet they are typically driven by separate scoring mechanisms. When a parameter mask and data subset jointly determine restricted fine-tuning, this separation incurs redundant overhead and makes coordinated selection difficult. We cast parameter and data selection as two bilevel selection problems under a common validation objective and derive a shared local response-surrogate scoring rule. Under first- and second-order validation-improvement approximations, parameter importance and data utility emerge as column-wise and row-wise aggregations of a single gradient interaction matrix, yielding a closed-form row-column correspondence for co-extracting both signals. Building on this structure, we propose DualSFT (Dual-Selection Fine-Tuning), a one-shot dual-scoring algorithm that produces a parameter mask and data subset from shared gradient statistics. On 3B-9B LLMs, single-axis DualSFT variants strengthen target-task performance and stability-plasticity trade-offs within their comparison groups, while full DualSFT yields a more favorable joint-constrained trade-off than sequential hybrid baselines under matched budgets.
Xinrui Chen, Liu Yang, Ou Wu
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences · College of Intelligence and Computing, Tianjin University
Fine-tuning large language models (LLMs) is compute-intensive and error-prone: model performance depends sensitively on data quality and hyperparameter choices, and naïve runs can even degrade model performance. This raises a practical question:can we predict fine-tuning performance before committing to a full training run? We present TUNEAHEAD, a lightweight framework for pre-hoc prediction of fine-tuning performance. TUNEAHEAD encodes each candidate run as a meta-feature vector that combines static dataset descriptors with dynamic probe features from a short standardized probe. A predictor maps these features to performance estimates, while SHAP-based attributions provide interpretable diagnostics that reveal which specific features drive the prediction. Across 1,300+ fine-tuning runs on Qwen2.5-7B-Instruct, TUNEAHEAD consistently outperforms strong baselines such as Early-Stop Extrapolation and ProxyLM. On a held-out test set of 370 runs, TUNEAHEAD achieves an RMSE of 1.47 percentage points and places 95.1% of predictions within +3/-3 percentage points of the true score. These accurate continuous predictions support practical go/no-go screening policies that can reduce unnecessary full fine-tuning while retaining most promising runs.
Yuxiang Luo, Haonan Long, Chen Wang +6
The Hong Kong University of Science and Technology (Guangzhou) · Huawei Technologies Ltd.