← Back to explorer

Sharpening Tax in Post-Training

Type
paper
Venue
arXiv:2610.01509 (cs.AI), submitted 1 Oct 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Tests the 'sharpening hypothesis' — that RL post-training merely sharpens a base model's existing behaviors, improving single-shot accuracy at the cost of solution coverage — in the agentic domain (multi-turn tool use), where post-training is believed to matter most. Pre-trained base LLMs with a light, model-agnostic inference harness turn out to be capable agents: they lose at pass@1 but scale more steeply with pass@k and often overtake their post-trained counterparts given a sufficient test-time budget. The mechanism: post-training pushes tasks toward two extremes (always solved or never solved), improving sampling efficiency and consistency at the cost of solution coverage. The paper proposes Sharpening Tax, a diagnostic metric quantifying the loss in test-time scalability after post-training (area between the pass@k curve and its pass@K ceiling, normalized; tax = base scalability minus post-trained scalability). Across 14 base/post-trained pairs from four model families on three agentic benchmarks (42 settings), the tax is prevalent (36 of 42 pay a positive tax at k=128), can be estimated cheaply from few rollouts (tax from 8 rollouts predicts tax at 32, Spearman rho = 0.85), and correlates with other metrics. Finally, posterior-tempered group sampling (PTGS) — a plug-and-play Bayesian sampler that adapts sampling temperature per prompt via a Beta posterior over difficulty — pays a smaller tax during RL training in two agentic environments, solving more tasks under repeated sampling while also improving pass@1.

Keywords

sharpening tax · post-training · RL · pass@k · distribution sharpening · PTGS · posterior-tempered group sampling · agentic benchmarks · WebShop · BFCL · ACEBench

Topics

post-training, RL, distribution sharpening, pass@k, agents, test-time scaling, sampling

Research notes

  • Discovery: @Changdae_Oh X thread (11 parts) 2026-10-01 (https://x.com/Changdae_Oh/status/2105852436602626399)
  • Project page: https://changdaeoh.github.io/sharpening-tax/
  • Code: https://github.com/changdaeoh/sharpening-tax
  • Teams: Meta Superintelligence Labs, UW-Madison, NYU, Stanford
  • Key thread details: with a simple model-agnostic harness, base models win at pass@1 for post-trained models, but base + light harness overtakes in 11 of 12 model-benchmark settings with more tries
  • Scale dependence (Gemma 4 on WebShop): tries for base to overtake its post-trained twin — never (4B), 16, 11, 3 (31B); sharpening lifts small models' floor but lowers large models' ceiling
  • Mechanism detail (Gemma 4 31B, WebShop, 128 trials/task): sometimes-solved 88% -> 30%, always-solved 0% -> 26%, never-solved 12% -> 44% after post-training
  • Tax definition: Tax = scalability of base - scalability of post-trained (normalized area between pass@k curve and pass@K ceiling); positive tax = post-training made repeated sampling less useful; 36 of 42 settings positive at k=128; cheap to forecast (8-rollout tax predicts 32-rollout tax, rho = 0.85)
  • PTGS: tracks each prompt's difficulty with a Beta posterior during RL — heat hard prompts, cool easy ones; plug-and-play for GRPO/PPO etc., no change to RL update, no extra compute; training Qwen2.5-7B-Instruct on Sokoban & FrozenLake improves both pass@1 and pass@128 and lowers the tax in all 4 settings
  • Benchmarks in heatmap: BFCL v4, WebShop, ACEBench across Gemma 4, Ministral 3, Qwen 2.5, Qwen 3.5
  • Takeaway: post-training mostly changes how reliably a model solves tasks, not which tasks it can solve
  • License: CC BY 4.0