Language models are increasingly used to sample from a specified distribution, for instance, to simulate survey respondents or generate synthetic data. Instruction-tuned models can state such a distribution correctly and still fail to sample from it. Prompting and changes to decoding reduce this mismatch only partly, which motivates training with policy optimization. Group relative policy optimization (GRPO) is a natural fit for this problem because it already samples a group of rollouts per prompt, and the group's empirical distribution can be compared with the target. However, scoring the group as a whole gives every rollout the same reward. Group-relative centering then sets all advantages to zero, and the model receives no learning signal. To give each rollout its own signal, we introduce the witness advantage, a per-rollout advantage derived from maximum mean discrepancy (MMD). It trains a model to match a target distribution over a finite set of outcomes. The MMD between the model's distribution and the target has a witness function that measures how over- or under-produced each outcome is. Each rollout's advantage estimates the negative witness at its outcome, so a rollout is rewarded for an outcome the group under-produces and penalized for one it over-produces. The witness advantage is computed in closed form from the group's outcome counts, and we use it as the reward in GRPO. On unseen target distributions, training with the witness advantage substantially reduces the total variation distance to the target while largely preserving the model's general capabilities.
Figures & tables
sampling error
capability
excess TV
MMLU
IFEval
untrained model
0.589
0.584
0.512
verbalized sampling
0.518
0.584
0.512
oracle temperature
0.478
0.584
0.512
supervised cross-entropy
0.406
0.482
0.281
witness advantage (ours)
0.376
0.583
0.528
Table 1: On the 100 targets from five distribution families never seen in training, the witness advantage has the lowest median excess TV and keeps general capabilities. The oracle temperature is tuned on each target.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
a (3 copies)
b
c
⊥
target probability q(x)
1/2
3/10
1/5
0
witness advantage Ai (ours)
1/5
3/5
2/5
0
full-group variant A^i
0
4/15
1/15
−1/3
Appendix
Table 2: The witness advantage and its full-group variant on a group of six rollouts with outcomes (a,a,a,b,c,⊥) , in exact fractions.
setting
value
model
Qwen2.5-1.5B-Instruct ( Qwen et al., 2024 ) , full fine-tuning, bfloat16 mixed precision
implementation
GRPO in TRL ( von Werra et al., 2020 ) , generation with vLLM ( Kwon et al., 2023 ) on the same GPU
group size G
64
prompts per step
4 ( 256 rollouts)
steps
600 , one gradient step per generation batch (on-policy)
Table 3: Settings shared by all GRPO runs. Exceptions are listed in the text.
untrained
cross-entropy
witness (ours)
sampling error
training targets
0.477
0.053
0.101
unseen parameters
0.483
0.094
0.161
unseen families
0.589
0.406
0.376
capability
MMLU
0.584
0.482
0.583
IFEval
0.512
0.281
0.528
Appendix
Table 4: Sampling error on the three synthetic evaluation sets and general capabilities. Supervised cross-entropy fits the training targets and unseen parameters more closely than the witness advantage but loses 10 MMLU and 23 IFEval points. Sampling error is the median excess TV on the 64 training and 16 unseen-parameter targets outside the Zipf family and on the 100 unseen-family targets. Capability is accuracy, with 95% bootstrap half-widths of 0.008 (MMLU) and 0.033 (IFEval).
training family
untrained
witness (ours)
group-scalar
sign witness
biased coin
0.361
0.034
0.440
0.021
geometric
0.372
0.050
0.309
0.028
binomial
0.516
0.176
0.511
0.255
Poisson
0.691
0.421
0.696
0.263
Zipf
0.563
0.024
0.447
0.523
Appendix
Table 5: Median excess TV per family. Top: training families after 600 steps, with the lowest value per family in bold. Bottom: unseen families and all 100 unseen-family targets pooled, for the untrained model and for the witness advantage trained on the original prompts or on prompts in the evaluation format. All values use n=500 draws per target. Figure 1 b uses the n assigned to each target, so its untrained values differ from those in this table.
median difference
witness advantage vs. group-scalar reward 16 unseen-parameter targets, no Zipf
Table 6: Paired per-target differences in TV on unseen targets outside the Zipf family. The top block uses the 16 unseen-parameter targets, and the bottom block adds 20 hypergeometric targets. Each difference is the TV of the second method minus the TV of the first, so a positive value favors the first. The superscript gives the larger side of the 95% bootstrap interval.
Table 7: Group-size sweep with 256 rollouts per step, seed 0. Each entry is the reduction in excess TV on the 100 unseen-family targets, in percent, and the superscript gives the larger side of the 95% paired bootstrap interval. All runs were trained on the original prompts and evaluated on the unseen-family prompts, so they compare with the 36% of the main run, not with the 24% of the retrain on the evaluation format. The G=64 witness run is separate from the main run.
method
excess TV
reduction
renormalization
0.600
−2%
seed conditioning
0.546
7%
verbalized sampling (best case)
0.518
12%
global temperature ( 2.76 )
0.514
13%
per-target oracle temperature
0.478
19%
witness advantage (ours)
0.376
36%
Appendix
Table 8: Training-free methods on the 100 unseen-family targets, applied to the untrained Qwen2.5-1.5B-Instruct, whose median excess TV is 0.589 . Reduction is relative to the untrained model. Both temperatures are chosen with knowledge of the targets.
training set
excess TV
reduction
one geometric target
0.397
33%
one family (16 geometric targets)
0.462
22%
five families (80 targets)
0.447
24%
fifteen families (240 targets)
0.394
33%
Appendix
Table 9: Median excess TV on the 100 unseen-family targets after training on sets of different variety, all with prompts in the evaluation format and the same compute. Reduction is relative to the untrained model ( 0.589 ).
opinion evaluation set
Qwen2.5-1.5B
Llama-3.1-8B
trained (ours)
reduction
urn draws
0.33
0.23
0.04
80 to 91%
GlobalOpinionQA
0.26
0.10
0.03
86 to 92%
NYTimes task
0.48
0.32
0.15
61 to 72%
implicit variant
0.31
0.21
0.11
56 to 69%
Appendix
Table 10: Stated opinion distributions and urn draws (top) and structured outputs (bottom), in median excess TV. Trained values are means over seeds, and reductions give the range over seeds: three for the opinion tasks, two for uniform sibling targets, and one for non-uniform sibling targets. The structured-output targets are evaluated on 60 test graphs.