Language models are increasingly used to sample from a specified distribution, for instance, to simulate survey respondents or generate synthetic data. Instruction-tuned models can state such a distribution correctly and still fail to sample from it. Prompting and changes to decoding reduce this mismatch only partly, which motivates training with policy optimization. Group relative policy optimization (GRPO) is a natural fit for this problem because it already samples a group of rollouts per prompt, and the group's empirical distribution can be compared with the target. However, scoring the group as a whole gives every rollout the same reward. Group-relative centering then sets all advantages to zero, and the model receives no learning signal. To give each rollout its own signal, we introduce the witness advantage, a per-rollout advantage derived from maximum mean discrepancy (MMD). It trains a model to match a target distribution over a finite set of outcomes. The MMD between the model's distribution and the target has a witness function that measures how over- or under-produced each outcome is. Each rollout's advantage estimates the negative witness at its outcome, so a rollout is rewarded for an outcome the group under-produces and penalized for one it over-produces. The witness advantage is computed in closed form from the group's outcome counts, and we use it as the reward in GRPO. On unseen target distributions, training with the witness advantage substantially reduces the total variation distance to the target while largely preserving the model's general capabilities.
Figures & tables
sampling error
capability
excess TV
MMLU
IFEval
untrained model
0.589
0.584
0.512
verbalized sampling
0.518
0.584
0.512
oracle temperature
0.478
0.584
0.512
supervised cross-entropy
0.406
0.482
0.281
witness advantage (ours)
0.376
0.583
0.528
Table 1: On the 100 targets from five distribution families never seen in training, the witness advantage has the lowest median excess TV and keeps general capabilities. The oracle temperature is tuned on each target.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
a (3 copies)
b
c
⊥
target probability q(x)
1/2
3/10
1/5
0
witness advantage Ai (ours)
1/5
3/5
2/5
0
full-group variant A^i
0
4/15
1/15
−1/3
Appendix
Table 2: The witness advantage and its full-group variant on a group of six rollouts with outcomes (a,a,a,b,c,⊥) , in exact fractions.
setting
value
model
Qwen2.5-1.5B-Instruct ( Qwen et al., 2024 ) , full fine-tuning, bfloat16 mixed precision
implementation
GRPO in TRL ( von Werra et al., 2020 ) , generation with vLLM ( Kwon et al., 2023 ) on the same GPU
group size G
64
prompts per step
4 ( 256 rollouts)
steps
600 , one gradient step per generation batch (on-policy)
Table 3: Settings shared by all GRPO runs. Exceptions are listed in the text.
untrained
cross-entropy
witness (ours)
sampling error
training targets
0.477
0.053
0.101
unseen parameters
0.483
0.094
0.161
unseen families
0.589
0.406
0.376
capability
MMLU
0.584
0.482
0.583
IFEval
0.512
0.281
0.528
Appendix
Table 4: Sampling error on the three synthetic evaluation sets and general capabilities. Supervised cross-entropy fits the training targets and unseen parameters more closely than the witness advantage but loses 10 MMLU and 23 IFEval points. Sampling error is the median excess TV on the 64 training and 16 unseen-parameter targets outside the Zipf family and on the 100 unseen-family targets. Capability is accuracy, with 95% bootstrap half-widths of 0.008 (MMLU) and 0.033 (IFEval).
training family
untrained
witness (ours)
group-scalar
sign witness
biased coin
0.361
0.034
0.440
0.021
geometric
0.372
0.050
0.309
0.028
binomial
0.516
0.176
0.511
0.255
Poisson
0.691
0.421
0.696
0.263
Zipf
0.563
0.024
0.447
0.523
Appendix
Table 5: Median excess TV per family. Top: training families after 600 steps, with the lowest value per family in bold. Bottom: unseen families and all 100 unseen-family targets pooled, for the untrained model and for the witness advantage trained on the original prompts or on prompts in the evaluation format. All values use n=500 draws per target. Figure 1 b uses the n assigned to each target, so its untrained values differ from those in this table.
median difference
witness advantage vs. group-scalar reward 16 unseen-parameter targets, no Zipf
Table 6: Paired per-target differences in TV on unseen targets outside the Zipf family. The top block uses the 16 unseen-parameter targets, and the bottom block adds 20 hypergeometric targets. Each difference is the TV of the second method minus the TV of the first, so a positive value favors the first. The superscript gives the larger side of the 95% bootstrap interval.
Table 7: Group-size sweep with 256 rollouts per step, seed 0. Each entry is the reduction in excess TV on the 100 unseen-family targets, in percent, and the superscript gives the larger side of the 95% paired bootstrap interval. All runs were trained on the original prompts and evaluated on the unseen-family prompts, so they compare with the 36% of the main run, not with the 24% of the retrain on the evaluation format. The G=64 witness run is separate from the main run.
method
excess TV
reduction
renormalization
0.600
−2%
seed conditioning
0.546
7%
verbalized sampling (best case)
0.518
12%
global temperature ( 2.76 )
0.514
13%
per-target oracle temperature
0.478
19%
witness advantage (ours)
0.376
36%
Appendix
Table 8: Training-free methods on the 100 unseen-family targets, applied to the untrained Qwen2.5-1.5B-Instruct, whose median excess TV is 0.589 . Reduction is relative to the untrained model. Both temperatures are chosen with knowledge of the targets.
training set
excess TV
reduction
one geometric target
0.397
33%
one family (16 geometric targets)
0.462
22%
five families (80 targets)
0.447
24%
fifteen families (240 targets)
0.394
33%
Appendix
Table 9: Median excess TV on the 100 unseen-family targets after training on sets of different variety, all with prompts in the evaluation format and the same compute. Reduction is relative to the untrained model ( 0.589 ).
opinion evaluation set
Qwen2.5-1.5B
Llama-3.1-8B
trained (ours)
reduction
urn draws
0.33
0.23
0.04
80 to 91%
GlobalOpinionQA
0.26
0.10
0.03
86 to 92%
NYTimes task
0.48
0.32
0.15
61 to 72%
implicit variant
0.31
0.21
0.11
56 to 69%
Appendix
Table 10: Stated opinion distributions and urn draws (top) and structured outputs (bottom), in median excess TV. Trained values are means over seeds, and reductions give the range over seeds: three for the opinion tasks, two for uniform sibling targets, and one for non-uniform sibling targets. The structured-output targets are evaluated on 60 test graphs.
Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from distributions, they collapse to a single output. The same persona on the same question returns the same answer on more than half of items in a public-opinion benchmark, and the model's internal probabilities concentrate on a single option. The failure is associated with, and amplified by, instruction-targeted post-training: instruction-tuned models are worse than their own bases in every family we can compare, the gap widens at each successive post-training stage and with the size of the tuning update, and continued pretraining on non-instruction tokens leaves it unchanged. Yet the knowledge survives: the same model that cannot sample from a distribution can describe it accurately in a single call. We call this gap the KNOWS/DOES split. Exploiting the split, a single call that asks the model to describe the response distribution more than halves the error against human survey data compared to persona aggregation. When per-persona outputs are required, we propose Prompt-Perturbed Argyle (PPA), which reduces the same error by 21%, spreading each persona's answers to mirror real population differences at no added cost.
Chaemin Jang, Dongman Lee, Jihee Kim
School of Computing, KAIST · School of Business and Technology Management, KAIST
GRPO-style RLVR trains reasoning models from multiple on-policy attempts per prompt, but typically uses these attempts only through terminal rewards. We show that a mixed group contains a richer process signal: a correct completion is a self-generated witness of how the current policy can solve the problem, while a wrong completion provides on-policy prefixes where the policy needs correction. We introduce \emph{Self-Supervised On-Policy Distillation} (SSOPD), which distills a teacher distribution conditioned on the shortest correct completion into prefixes of the longest wrong completion. This converts intra-group correct--wrong contrast into dense process supervision without external solution traces. A stopping-time view motivates the shortest-correct / longest-wrong rule as a finite-group approximation to editing persistent failures toward fast-success actions, and a prompt-level frontier weight concentrates the auxiliary loss where correct and wrong branches coexist. Across AIME 2024, AIME 2025, and HMMT 2025, SSOPD improves over GRPO in all nine model-benchmark settings. On Qwen3-8B, it reaches a macro Avg@12 of 65.6, outperforming GRPO by 1.6 points and the solution-conditioned OPSD baseline by 0.8 points. Code will be released at https://github.com/tzq1999/SSOPD.
Three of the most popular methods for training language models to reason look like three different tricks. They are not. All three adjust a single number: standard deviation, reflecting how much a prompt's sampled answers disagree. When such a model is trained, it answers each problem many times, and an automatic checker marks every answer right or wrong. The standard deviation of those marks measures the disagreement: largest when the answers split evenly between right and wrong, and zero when they all agree. Group Relative Policy Optimization (GRPO) divides by this number, GRPO Done Right (Dr. GRPO) drops the division, and Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) discards the groups where it is zero. Each is presented as its own fix, yet this paper proves they are three settings of one dial. That dial is not cosmetic: for right-or-wrong rewards, the disagreement is exactly the size of the training update, the group-standard-deviation identity. A split group teaches the most, while a unanimous group teaches nothing and falls silent. The same result says which problems deserve the most weight and how many tries each one needs. This paper confirms the intuition on a large real difficulty dataset (Big-Math) and in a controlled training run. What looks like a harmless normalization step is the dial that decides where learning happens and how strongly.