FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training
Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Abstract
Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matched learner study reports quality 0.6329 for the fixed selector, compared with 0.5482 for uniform and 0.6037 for format-feedback under 128 updates. Metadata-only cross-fitted calibration reaches over eight seeds (median 0.6481; paired 95% interval ) at 63,276 target-run tokens; its recorded full cost is 189,642 tokens and 3.48 GPU-hours including calibration. The completed FAER-UTILITY row reaches 0.6624 at 62,844 target-run tokens and 4.26 GPU-hours. Format-feedback selects records with correctness 0.6953, compared with 0.3594 for the fixed selector, despite the different downstream ranking. The completed comparison surfaces report the learner-aware ablation, same-seed gap, policy-optimization rows, and strict zero-shot transfer.
Figures & tables
| Stage | Population / budget | Replication |
|---|---|---|
| Fixed cache A | 64 train / 128 test; 32 selected | 8 selection seeds |
| Fixed cache B | 256 train / 1,319 test; 32 selected | 8 selection seeds |
| Measured age | 192 family-A + 1,575 family-B records | independent producers |
| Matched learner | 256 / 128 / 256 train/cal/test tasks; 640 rows per seed | 3 independent seeds; 1,920 rows |
| Selector | Quality | Target tokens | Total tokens | GPU-h | 95% CI |
|---|---|---|---|---|---|
| Uniform | 0.5482 | 65,214 | 65,214 | 0.99 | reference |
| Format-feedback | 0.6037 | 64,982 | 64,982 | 0.98 | |
| FAER-Fixed | 0.6329 | 63,771 | 63,771 | 0.96 | |
| Token-efficiency | 0.6186 | 63,428 | 63,428 | 0.95 | |
| FAER-Utility-Meta | 0.6476 | 63,276 | 189,642 | 3.48 | |
| FAER-Utility | 0.6624 | 62,844 | 189,210 | 4.26 |
| Arm set | Blocks | ||||
|---|---|---|---|---|---|
| Four arms | 8 | 0.1250 | 0.4500 | 0.4375 | 0.0294 |
| 95% block CI | 8 |
| Selector | Quality | Target tokens | Paired 95% CI |
|---|---|---|---|
| FAER-Fixed | 0.6329 | 63,771 | |
| FAER-Utility-Meta | 0.6476 | 63,276 | |
| Alignment-only | 0.6418 | 63,086 | |
| FAER-Utility | 0.6624 | 62,844 | reference |
| Shuffled alignment | 0.6484 | 63,171 | |
| Random calibration gradient | 0.6447 | 63,263 |
Appendix figures & tables47 assets
Supplementary material from the paper’s appendix.
Appendix
| Field | Role | Audit treatment |
|---|---|---|
| Trajectory ID | sampling and label join key | included in trace; duplicate IDs rejected |
| Format status | observed feedback priority | serialized and taint-tested |
| Verifier confidence | observed reliability priority | producer key recorded; generation lineage separate |
| Model confidence | observed confidence/surprise | serialized and taint-tested |
| Token count | length-normalized priority and telemetry | producer field and units recorded |
| Target answer | offline GSM8K reference | excluded from selector projection |
| Field | Recorded value |
|---|---|
| Producer | build_cache.py::produce ; source prefix 4c7a1e9d |
| Inputs | prompt/model-output; 0 pre-freeze target accesses |
| Order | generate extract verify serialize freeze oracle join |
| Runtime | Qwen2.5-1.5B-Instruct; verifier-v1.2; environment prefix b91d63af |
| Lineage | 192/192 records; 0/192 process-isolation violations |
| Selector | 0 oracle-key hits; 1,344/1,344 invariant mutations |
| Field | Setting / result |
|---|---|
| Model | Qwen2.5-1.5B-Instruct |
| Optimizer / precision | AdamW / bf16 |
| Learning rate | |
| Update / token / verifier caps | 128 / 65,536 generated tokens / 4,096 calls |
| Stopping rule | 128 updates or generated-token cap |
| Independent seeds | 13, 17, 23 |
| Policy | Quality | Tokens | p50/p95 (s) | 95% interval |
|---|---|---|---|---|
| Reliability | 0.5948 | 64,284 | 0.821/1.276 | |
| Uncertainty-PER | 0.5817 | 64,736 | 0.835/1.309 |
| Package | Cache / task identity | Comparison axis |
|---|---|---|
| (a) Learner update | family A; Qwen/GSM8K | replay policy at shared training budget |
| (b) Freshness + family | families A and B; Qwen/GSM8K | measured age and producer family |
| (c) End-to-end | second family release; Qwen2.5-1.5B-Instruct; GSM8K | task quality and generated-token cost |
| Work | Replay unit / priority | Freshness handling |
|---|---|---|
| FreshPER | stored experience with PER priority multiplied by exponential age decay | explicit freshness/age decay |
| TBA | central replay-buffer data sampled by reward or recency | explicit recency prioritization |
| RLEP | verified successful trajectories mixed with newly generated rollouts | two-phase replay; no explicit per-row age priority |
| DOTS/RR | difficulty-targeted online selection plus reuse of recent rollouts | recent-rollout replay |
| Prioritized Replay for RL Post-training | problem-level priority from empirical success statistics | periodic retesting |
| Rollout-Level Advantage-PER | individual rollouts prioritized by advantage magnitude | age eviction and fresh-anchored composition |
| Work | Update treatment | Verifier / target timing |
|---|---|---|
| FreshPER | PER replay with exponential age decay; no separate importance-ratio correction | reward/task signal enters the RL update; no FAER-style post-freeze target join |
| TBA | Trajectory Balance objective on off-policy data | reward enters the TB update; no post-freeze selector/evaluator boundary |
| RLEP | fresh rollouts mixed with replayed verified successes at each update | verification determines replay eligibility before optimization |
| DOTS/RR | recent-rollout reuse under GRPO-style policy updates | reward/difficulty statistics are training-time signals; no post-freeze oracle split |
| Prioritized Replay for RL Post-training | adaptive problem resampling under GRPO; no trajectory-level importance correction | empirical success statistics are training-time priority signals |
| Rollout-Level Advantage-PER | fresh-anchored GRPO batches plus age-bounded replay | verifiable reward/advantage enters priority before replay update |
| field | type and role | acceptance record |
|---|---|---|
| trajectory ID | stable string; cache join key | unique within split; train/test disjoint |
| trajectory text | generated sequence; replay unit | digest-bound byte sequence |
| format status | binary observed feedback | producer key and value domain |
| verifier confidence | observed score | producer key, version, and units |
| model confidence | observed score | producer key, version, and units |
| token count | nonnegative integer telemetry | tokenizer identity and count rule |
| gate | check | artifact field |
|---|---|---|
| identity | source, model, data, split, and row-count hashes | manifest digest |
| schema | types, required keys, unique IDs, and disjoint splits | validation report |
| taint | target-field mutation preserves projection and priorities | 1,344/1,344 result |
| trace | seed, ordered IDs, priorities, and replay digest agree | trace manifest |
| oracle | post-freeze answer join and denominator are complete | evaluator digest |
| learner | matched model, optimizer, budget, and task IDs | 18/18 arm-seed runs accepted; checkpoint/metrics digests complete |
| Policy | Train | Test | test–train |
|---|---|---|---|
| Uniform | 13/32 (0.4062) | 12/32 (0.3750) | |
| Format-feedback | 15/32 (0.4688) | 22/32 (0.6875) | |
| Reliability | 14/32 (0.4375) | 22/32 (0.6875) | |
| Surprise | 1/32 (0.0312) | 5/32 (0.1562) | |
| Uncertainty | 6/32 (0.1875) | 4/32 (0.1250) | |
| Feedback-aware | 12/32 (0.3750) | 14/32 (0.4375) |
| Contrast | correct | tokens | 95% interval |
|---|---|---|---|
| Format-feedback uniform | |||
| Feedback-aware format-feedback | |||
| Token-efficiency format-feedback | |||
| Reliability format-feedback |
| Family/policy | Coverage | Correct | Age (h) | Digest |
|---|---|---|---|---|
| A / format-feedback | 192/192 (100%) | 0.7016 | 23.7 h | a7c91e42 |
| A / age-weighted | 192/192 (100%) | 0.7312 | 9.6 h | 5f2bd8c1 |
| B / format-feedback | 1,575/1,575 (100%) | 0.5849 | 31.4 h | c83e14aa |
| B / age-weighted | 1,575/1,575 (100%) | 0.6617 | 11.8 h | 2de709f6 |
| Policy | Quality | Tokens | p50/p95 (s) |
|---|---|---|---|
| Uniform | 0.5482 | 65,214 | 0.842/1.327 s |
| Format-feedback | 0.6037 | 64,982 | 0.836/1.305 s |
| Feedback-aware | 0.6329 | 63,771 | 0.829/1.289 s |
| Token-efficiency | 0.6186 | 63,428 | 0.806/1.251 s |
| Policy | 95% interval | Checkpoint |
|---|---|---|
| Uniform | reference | step-128@61c8a94e |
| Format-feedback | step-128@73ea20bd | |
| Feedback-aware | step-128@c2f47a81 | |
| Token-efficiency | step-128@95b16f3c |
| Arm | Tokens | p50/p95 (s) | Digest |
|---|---|---|---|
| FAER-product | 755 | 1.24/1.61 | 8d31f2c4 |
| Uniform | 751 | 1.21/1.58 | 54b92a17 |
| Format-feedback | 744 | 1.19/1.55 | c7e6409b |
| Reliability | 746 | 1.20/1.56 | 9a4cf6e2 |
| Token-efficiency | 739 | 1.18/1.53 | f13e8b75 |
| Check | Result |
|---|---|
| Answer extraction | 191/192 (99.48%); 1 unresolved |
| Adjudicated agreement | 95/96 (98.96%); 1 changed label |
| Numeric forms | 47/48 exact; 48/48 tolerant-equivalent |
| Cache reconstruction | 192/192 row digests match |
| Trace reconstruction | 56/56 ordered traces match |
| Table regeneration | 176/176 numeric cells match |
| setting | policy | decay | correct | SD | Jaccard |
|---|---|---|---|---|---|
| no decay | freshness | 0.0000 | 0.3047 | 0.0534 | 0.1377 |
| no decay | feedback-aware | 0.0000 | 0.3359 | 0.0644 | 0.1976 |
| no decay | format-feedback | 0.0000 | 0.7031 | 0.0180 | 0.6998 |
| mild recent | freshness | 0.0200 | 0.2891 | 0.0534 | 0.2312 |
| mild recent | feedback-aware | 0.0200 | 0.3125 | 0.0807 | 0.2399 |
| mild recent | format-feedback | 0.0200 | 0.6562 | 0.0000 | 0.7457 |
| Policy | Seed 13 | Mean SD | Min–max | Jaccard |
|---|---|---|---|---|
| Uniform | 12/32 (0.3750) | 0.2891 0.0779 | 0.1875–0.4062 | 0.1368 |
| Format-feedback | 22/32 (0.6875) | 0.6953 0.0277 | 0.6562–0.7500 | 0.6917 |
| Reliability | 22/32 (0.6875) | 0.4922 0.1053 | 0.3750–0.6875 | 0.2220 |
| Surprise | 5/32 (0.1562) | 0.0781 0.0472 | 0.0312–0.1562 | 0.2210 |
| Uncertainty | 4/32 (0.1250) | 0.2070 0.0471 | 0.1250–0.2812 | 0.1771 |
| Feedback-aware | 14/32 (0.4375) | 0.3594 0.0625 | 0.2812–0.4375 | 0.2162 |
| Policy | Family A | Family B |
|---|---|---|
| Uniform | 0.2891 | 0.3516 |
| Format-feedback | 0.6953 | 0.5781 |
| Reliability | 0.4922 | 0.4258 |
| Uncertainty | 0.2070 | 0.2344 |
| Feedback-aware | 0.3594 | 0.3594 |
| Token-efficiency | 0.6328 | 0.6523 |
| priority | mean | sample SD | min | max |
|---|---|---|---|---|
| uniform | 0.2891 | 0.0779 | 0.1875 | 0.4062 |
| format-feedback | 0.6953 | 0.0277 | 0.6562 | 0.7500 |
| reliability | 0.4922 | 0.1053 | 0.3750 | 0.6875 |
| surprise | 0.0781 | 0.0472 | 0.0312 | 0.1562 |
| uncertainty | 0.2070 | 0.0471 | 0.1250 | 0.2812 |
| feedback-aware | 0.3594 | 0.0625 | 0.2812 | 0.4375 |
| priority | train | test | test–train |
|---|---|---|---|
| uniform | 0.4062 | 0.3750 | |
| format-feedback | 0.4688 | 0.6875 | |
| reliability | 0.4375 | 0.6875 | |
| surprise | 0.0312 | 0.1562 | |
| uncertainty | 0.1875 | 0.1250 | |
| feedback-aware | 0.3750 | 0.4375 |
| priority | correct | format | verifier | tokens |
|---|---|---|---|---|
| uniform | ||||
| format-feedback | ||||
| reliability | ||||
| surprise | ||||
| uncertainty | ||||
| feedback-aware |
| Package | Selected-set quality | Learner loss | Task quality |
|---|---|---|---|
| (a) Learner | |||
| (b) Age/family | |||
| (c) End-to-end |
| Package | Generated tokens | p50/p95 (s) | Missing age |
|---|---|---|---|
| (a) Learner | 0.829/1.289 | 0/1,920 | |
| (b) Age/family | 0.812/1.263 | 0/1,767 | |
| (c) End-to-end | 0.791/1.221 | 0/1,920 |
| Priority | Task quality | Tokens | Paired 95% interval |
|---|---|---|---|
| Full product | 0.6329 | 63,771 | reference |
| Remove uncertainty | 0.6164 | 64,093 | |
| Remove surprise | 0.6088 | 64,261 | |
| Monotone verifier factor | 0.6205 | 63,984 | |
| Calibrated verifier | 0.6261 | 63,856 | |
| Remove format feedback | 0.6109 | 64,182 |
| Axis | Setting | Task quality | Generated tokens |
|---|---|---|---|
| Update horizon | 32 updates | 0.5787 | 16,117 |
| Update horizon | 64 updates | 0.6089 | 31,986 |
| Update horizon | 128 updates, FAER | 0.6329 | 63,771 |
| Measured-age decay | 0.6214 | 63,486 | |
| Measured-age decay | selected | 0.6418 | 63,204 |
| Measured-age decay | 0.6306 | 62,917 |
| Control | Trace Jaccard | Task quality | Paired interval |
|---|---|---|---|
| Original fields | 1.0000 | 0.6329 | reference |
| Verifier permutation | 0.3614 | 0.5946 | |
| Age permutation within run | 1.0000 | 0.6329 | |
| Format-status permutation | 0.4278 | 0.6072 |
| Selector | Final quality | Reward | Generated tokens | KL |
|---|---|---|---|---|
| Uniform | 0.5594 | 0.5217 | 65,188 | 0.0379 |
| Format-feedback replay | 0.6031 | 0.5758 | 64,736 | 0.0436 |
| FreshPER | 0.6226 | 0.6019 | 64,127 | 0.0418 |
| FAER-Fixed | 0.6387 | 0.6194 | 63,692 | 0.0427 |
| FAER-Utility | 0.6575 | 0.6426 | 62,961 | 0.0443 |
| Seeds | Policy | Mean SD | Median | 95% CI | Tokens |
|---|---|---|---|---|---|
| 5 | uniform | 0.5508 | reference | 65,102 | |
| 5 | feedback-aware | 0.6332 | 63,846 | ||
| 5 | utility-calibrated | 0.6497 | 63,318 | ||
| 8 | uniform | 0.5486 | reference | 65,241 | |
| 8 | feedback-aware | 0.6305 | 63,802 | ||
| 8 | utility-calibrated | 0.6481 | 63,276 |
| Policy | Replay unit | Priority / fitting | Quality | Tokens |
|---|---|---|---|---|
| FreshPER adaptation | complete trajectory | PER with age decay | 0.6242 | 63,688 |
| Uniform | complete trajectory | equal weight | 0.5482 | 65,214 |
| Format-feedback | complete trajectory | observed format signal | 0.6037 | 64,982 |
| Fixed FAER | complete trajectory | fixed product | 0.6329 | 63,771 |
| Surprise-only | complete trajectory | 0.5893 | 65,407 | |
| Age-only | complete trajectory | 0.6418 | 63,204 |
| Task | Model | Fixed FAER | Utility-calibrated FAER |
|---|---|---|---|
| GSM8K | Qwen2.5-1.5B-Instruct | 0.6329 / 63,771 | 0.6489 / 62,914 |
| MATH-500 | Qwen2.5-1.5B-Instruct | 0.3816 / 65,318 | 0.3954 / 65,104 |
| SVAMP | Qwen2.5-1.5B-Instruct | 0.7137 / 58,664 | 0.7284 / 58,203 |
| GSM8K | Qwen2.5-7B-Instruct | 0.7043 / 62,087 | 0.7168 / 61,642 |
| Target task | Target model | Fixed | Utility-Meta | Utility |
|---|---|---|---|---|
| MATH-500 | Qwen2.5-1.5B-Instruct | 0.3816 | 0.3954 | 0.4047 |
| SVAMP | Qwen2.5-1.5B-Instruct | 0.7137 | 0.7284 | 0.7369 |
| GSM8K | Qwen2.5-7B-Instruct | 0.7043 | 0.7168 | 0.7256 |
| MBPP | Qwen2.5-1.5B-Instruct | 0.4128 | 0.4271 | 0.4369 |
| Treatment | Transform | Brier | Correct | Jaccard |
|---|---|---|---|---|
| Fixed product | 0.1728 | 0.3594 | 1.0000 | |
| Remove uncertainty | 0.1609 | 0.4219 | 0.4386 | |
| Reported monotone | 0.1587 | 0.4531 | 0.4019 | |
| Monotone quadratic | 0.1616 | 0.4375 | 0.3864 | |
| Calibrated verifier | 0.1472 | 0.4062 | 0.5173 |
| Selector | Quality | Tokens | GPU-h | 95% CI |
|---|---|---|---|---|
| FAER-Utility-Meta (cross-fitted control) | 0.6494 | 189,642 | 3.48 | |
| Random weights | 0.5893 | 251,208 | 3.32 | |
| Permuted fitted weights | 0.6048 | 250,734 | 3.36 | |
| Entropy-matched fixed FAER | 0.6336 | 191,485 | 2.72 | |
| Correctness-calibrated selector | 0.6124 | 188,906 | 2.68 |
| Selector | Diversity | Duplicates | Quality | Tokens |
|---|---|---|---|---|
| Format-feedback | 0.4127 | 0.1562 | 0.6037 | 64,982 |
| Fixed FAER | 0.5274 | 0.0625 | 0.6329 | 63,771 |
| Diversity-matched format-feedback | 0.5196 | 0.0703 | 0.6174 | 64,336 |
| Estimator | Spearman | Kendall | Sign acc. | Top-32 regret |
|---|---|---|---|---|
| Gradient cosine | 0.4176 | 0.2918 | 0.7064 | 0.0473 |
| Raw gradient dot product | 0.5062 | 0.3617 | 0.7441 | 0.0386 |
| Virtual update, first-order | 0.6827 | 0.4984 | 0.8019 | 0.0214 |
| Virtual update + norm correction | 0.7489 | 0.5681 | 0.8376 | 0.0127 |
| Shuffled virtual utility | 0.0216 | 0.0149 | 0.5078 | 0.0946 |
| Random optimizer-state control | 0.0693 | 0.0481 | 0.5297 | 0.0868 |
| Utility estimator | Quality | Target tokens | Paired 95% CI |
|---|---|---|---|
| Cosine alignment | 0.6418 | 63,086 | |
| Raw gradient dot product | 0.6473 | 63,014 | |
| Virtual SGD update | 0.6511 | 62,967 | |
| Virtual AdamW update, first-order | 0.6572 | 62,901 | |
| Virtual AdamW + norm correction | 0.6624 | 62,844 | reference |
| Shuffled virtual utility | 0.6484 | 63,171 |
| Selector | Target-matched quality | Total-matched quality | Total tokens | GPU-h |
|---|---|---|---|---|
| Uniform | 0.5482 | 0.6173 | 189,176 | 4.25 |
| Format-feedback | 0.6037 | 0.6358 | 189,094 | 4.24 |
| FreshPER | 0.6242 | 0.6446 | 189,128 | 4.23 |
| FAER-Fixed | 0.6329 | 0.6505 | 189,163 | 4.24 |
| FAER-Utility | 0.6624 | 0.6624 | 189,210 | 4.26 |
| Selected | Passes | Uniform | Format-feedback | FAER-UTILITY |
|---|---|---|---|---|
| 16 | 4 | 0.5239 | 0.5796 | 0.6371 |
| 32 | 1 | 0.4981 | 0.5492 | 0.6031 |
| 32 | 2 | 0.5276 | 0.5834 | 0.6365 |
| 32 | 4 | 0.5482 | 0.6037 | 0.6624 |
| 64 | 2 | 0.5563 | 0.6111 | 0.6581 |
| 128 | 1 | 0.5604 | 0.6066 | 0.6518 |
| Experiment | Recorded outcome |
|---|---|
| Fixed-cache selection | family A: format-feedback mean 0.6953, FAER-Fixed 0.3594; 8 draws |
| Paired inclusion | 128/128 marginal; 8,128/8,128 joint; four contrasts |
| Formula / index age | verifier peak at 0.75; four settings and four seeds |
| Measured age | 1,767/1,767 records; families A and B |
| Producer / selector | 192/192 lineage; 0 isolation violations; 1,344 invariant mutations |
| Evaluator | 191/192 extraction; 95/96 adjudicated agreement |
| Field | Value |
|---|---|
| Source commit / environment SHA-256 prefix | 197ec447 / b91d63af |
| Learner model and tokenizer | Qwen2.5-1.5B-Instruct, one checkpoint and tokenizer shared by all arms |
| Learner loss / target construction / token masking | masked autoregressive NLL / GSM8K target answer / assistant-response tokens only |
| 32-record selection to 128-update schedule / reuse | one complete trajectory per update / seeded cyclic order, four complete passes |
| Replay-probability correction in the learner | none; equal token loss weight |
| Decoder configuration / sampler implementation | greedy; do_sample=false , top_p=1.0 , max_new_tokens=512 / sequential weighted draw without replacement |
| stage | unit and bound artifact | reported metric |
|---|---|---|
| cache | split manifest and row digest | row count, ID disjointness, field coverage |
| projection | one row’s observed-field serialization | oracle-key hits, taint invariance |
| priority | row-policy pair and formula version | priority value, support floor, rank |
| sampling | seed and ordered without-replacement draw | selected IDs, trace digest, Jaccard |
| oracle join | selected ID and evaluator record | correct count, denominator, failure class |
| learner | arm, checkpoint, and sealed task IDs | quality, tokens, latency, paired interval |
| event | trigger and localization | resolution record |
|---|---|---|
| field | type/range/producer-key mismatch | row ID, key, expected domain, digest |
| taint | target mutation changes observed view or priority | mutation ID, field path, priority pair |
| trace | ordered draw or replay digest mismatch | seed, draw index, expected/actual ID |
| join | selected ID lacks evaluator correspondence | task ID, split, evaluator hash |
| label | extraction, normalization, or tolerance disagreement | adjudication pair and rule ID |
| update | checkpoint, budget, or sealed-task mismatch | arm, step, checkpoint, metric digest |