cs.AIOct 1, 2026

Sharpening Tax in Post-Training

Authors: Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov, Deren Lei, Yun He, Hoang Phan, Hangoo Kang, +2 more

Organizations: Meta Superintelligence Labs · University of Wisconsin–Madison · Work done at Meta · NYU · Stanford University

Abstract

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.

Figures & tables

Appendix figures & tables40 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

    Aug 10, 2026Ting Zhou, Zhenqing Ling, Daoyuan Chen +4Reasoning BenchmarkPost-Training

  2. Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

    Jun 24, 2026Changdae Oh, Wendi Li, Seongheon Park +3Reinforcement Learning Post-TrainingAgentic Reinforcement Learning

  3. LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents

    Jun 16, 2026Haoyang Fang, Wei Zhu, Boran Han +11Reinforcement Learning Post-TrainingLarge Language Model Training