Driven by recent advances in artificial intelligence, a growing literature has demonstrated the potential of using large language models (LLMs) as scalable surrogates to generate human-like responses. Two common approaches to improve the performance of LLMs include: fine-tuning, which aligns the LLM more closely with human responses, and rectification, which corrects biases in LLM outputs. In this paper, we develop a two-stage framework that combines fine-tuning and rectification, and optimally allocates limited labeled samples across the two stages. A key insight is that the conventional fine-tuning objective of minimizing mean squared prediction error is generally not aligned with the downstream rectification stage. For mean estimation, we propose to minimize the variance of the prediction errors; for general M-estimation, we propose to minimize a scalarized variance metric as the fine-tuning objective. Building on this insight, we leverage the scaling law of fine-tuning to optimally allocate the limited labeled human data between the fine-tuning and rectification stages. Our empirical analysis validates the fine-tuning scaling law and confirms that our proposed optimal allocation rule reliably identifies the optimal sample allocation. We demonstrate substantial efficiency gains in estimation and inference performance relative to fine-tuning or rectification alone, or to employing the conventional mean squared error objective within the fine-tuning then rectification framework. Such efficiency gains translate to significant cost savings for making reliable decisions.
Figures & tables
Review
This is a walk backward after the impressive 2012. Almost impenetrably black, the flavors converge around espresso and bitter chocolate, yet the tannins have a green edge. The wine simply feels flat in the mouth with no life to it.
Price
$28
Rating ( Y )
85
Review
This is one of the great Rieslings from the Wachau, a wonderful panoply of ripe, tropical fruit, pierced with flint, spice and minerality. It is rich and opulent, while never losing sight of the core tautness of a fine Riesling.
Price
$70
Rating ( Y )
95
Table 1 : Examples from the Wine Reviews dataset
Figure 1 : Empirical validation of the variance-based scaling law. The estimated parameters are a=11.06 , α=0.318 , and b=2.17 (adjusted R2=0.999 ); points are medians over ten training replicates per size with 95% error bars ( 1.96 standard errors over the replicates)
Figure 2 : Mean estimated 95% CI width of the FT+PPI estimator across sample allocations, for the variance-loss and MSE-loss fine-tuning objectives.
Figure 3 : Relative variance reduction of FT+PPI versus the Sample Mean , computed from the fitted scaling law.
n
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE ( ×10−2% )
1000
Sample Mean
10.062(0.432)
8.018(0.352)
9.060(0.397)
FT-Only
32.462(0.469)
31.308(0.496)
35.380(0.561)
PPI-Only
7.092(0.306)
5.621(0.250)
6.352(0.283)
FT+PPI (MSE Loss)
9.767(0.432)
7.653(0.351)
8.649(0.397)
FT+PPI (Var Loss)
7.155(0.261)
5.819(0.241)
6.575(0.272)
3000
Sample Mean
5.783(0.226)
4.600(0.203)
5.198(0.229)
Table 2 : Estimation accuracy across estimators under different labeled sample sizes n
n
FT+PPI (Var Loss) vs. Sample Mean
FT+PPI (MSE Loss) vs. Sample Mean
FT+PPI: Var vs. MSE Loss
1000
47.84(4.43)
7.83(14.47)
42.27(8.87)
3000
53.02(3.24)
31.22(4.54)
31.46(5.87)
5000
53.91(2.15)
37.05(3.14)
26.65(4.35)
10000
51.90(1.51)
38.36(1.92)
21.91(2.89)
Table 3 : Variance reductions (%) across estimators
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE (%)
OLS
2.436(0.100)
1.956(0.084)
0.997(0.043)
PPI-Only
1.679(0.066)
1.341(0.058)
0.684(0.030)
FT+PPI (MSE Loss)
1.420(0.055)
1.155(0.048)
0.589(0.024)
FT+PPI (Var Loss)
1.404(0.051)
1.145(0.047)
0.584(0.024)
Table 4 : Estimation accuracy for the price-rating regression coefficient at n=10,000
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE (%)
OLS
2.976(0.148)
2.394(0.125)
3.303(0.173)
FT+PPI (MSE Loss)
3.839(0.185)
3.063(0.164)
4.227(0.226)
FT+PPI (Var Loss)
3.259(0.165)
2.599(0.139)
3.586(0.192)
Table 5 : Controlled synthetic study: estimation accuracy for the census slope at n=1,200
Figure 4 : Bootstrap robustness of the scaling-law fit based on 5,000 -sample development pools. Light curves correspond to individual bootstrap replicates, the solid line denotes the median fit, and the shaded region indicates the central 95% robustness interval.
Quantity
Median
95% Robustness Interval
Scaling parameter a
14.265
[10.057,63.832]
Scaling parameter α
0.391
[0.251,0.531]
Variance floor b
2.338
[0.990,2.434]
Optimal allocation s∗/n
0.104
[0.094,0.107]
Table 6 : Robustness of the scaling-law fit
Figure 5 : Validation residual variance under variance-based and MSE-based fine-tuning, with the constant-predictor baseline ( Var(Y) ) for reference; current recipe (frozen text-embedding-3-large + MLP head).
n
RMSE (Cross-entropy) ( ×10−2 )
RMSE (Var Loss) ( ×10−2 )
1,000
10.038
7.873
3,000
4.303
2.567
5,000
3.037
2.816
10,000
2.471
2.514
Table 7 : Cross-entropy training versus the variance loss
Estimator
Residual variance
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE ( ×10−2% )
FT+PPI (MSE Loss)
3.626
2.187
1.743
1.969
FT+PPI (Var Loss)
3.205
1.951
1.580
1.786
Table 8 : LoRA fine-tuning: residual variance and PPI estimation accuracy
Design
Training Objective
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE ( ×10−2% )
Sample Mean
—
3.374(0.121)
2.766(0.112)
3.126(0.126)
PPI-Only
—
2.473(0.104)
1.991(0.085)
2.250(0.096)
FT+PPI split
MSE
2.381(0.094)
1.892(0.084)
2.138(0.094)
FT+PPI split
Variance
2.177(0.082)
1.752(0.075)
1.980(0.084)
Cross-fitting K=3
MSE
1.909(0.083)
1.509(0.068)
1.706(0.076)
Variance
1.734(0.081)
1.372(0.061)
1.551(0.069)
Table 9 : Cross-fitting on the main corpus: training objective at each fold count
n
Training Objective
PPI Variance
PPI++ Variance
Mean λ
1,000
MSE
9.38×10−3
7.58×10−3
0.61
Variance
5.17×10−3
5.06×10−3
0.97
10,000
MSE
6.11×10−4
5.42×10−4
0.82
Variance
4.79×10−4
4.65×10−4
0.95
Table 10 : Estimator variance under plain PPI and PPI++
Zero-Shot Model
Corr. with Human Ratings
Variance reduction w.r.t. Sample Mean
gpt-5.4-nano
0.741
52% – 54%
gpt-5.6-terra
0.860
72% – 74%
gpt-5.6-sol
0.863
73% – 74%
gpt-5.6-luna
0.885
77% – 78%
Table 11 : Zero-shot model quality and estimator efficiency under affine calibration
n
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE (%)
1,000
Sample Mean
3.286(0.328)
2.905(0.352)
0.694(0.084)
PPI-Only
2.293(0.341)
1.886(0.299)
0.451(0.071)
FT+PPI (MSE Loss)
3.728(0.445)
3.079(0.482)
0.736(0.115)
FT+PPI (Var Loss)
2.539(0.463)
2.041(0.347)
0.488(0.083)
2,000
Sample Mean
2.504(0.462)
1.923(0.368)
0.459(0.088)
PPI-Only
1.534(0.135)
1.384(0.152)
0.331(0.036)
Table 12 : Customer-review corpus: estimation accuracy across estimators
n
Estimator
RMSE ( ×10−2 )
MAE ( ×10−2 )
MAPE (%)
250
Sample Mean
8.603(1.166)
6.997(1.149)
1.881(0.309)
PPI-Only (prior wave)
5.251(0.591)
4.402(0.657)
1.183(0.177)
FT+PPI (MSE Loss)
8.452(1.055)
7.138(1.039)
1.919(0.279)
FT+PPI (Var Loss)
8.798(1.250)
6.841(1.269)
1.839(0.341)
500
Sample Mean
6.245(0.728)
5.309(0.755)
1.427(0.203)
PPI-Only (prior wave)
4.732(0.669)
3.878(0.622)
1.043(0.167)
Table 13 : Digital-twin survey panel: estimation accuracy across estimators
Component
Role
Value
Target-population size
census size
4,000
Cluster location zhi
center of the high-influence cluster
3.5
Cluster mass
share of the census in the cluster
15%
Cluster signal weight κ
outcome weight of the high-influence component
3.0
Outcome noise σ
homoskedastic noise scale
0.8
Nuisance dimensions
uninformative embedding coordinates
5
Table 14 : Generating parameters of the synthetic design