On the possibilities we keep—and the ones we lose
Sharpening Tax
in Post-Training
A better first try. A smaller world of solutions?
* Work done at Meta
01 / Background & motivation
New capabilities.
Or sharper probabilities?
The sharpening hypothesis: post-training amplifies behaviors already present in the base model, improving first-try accuracy while narrowing solution coverage.
The missing piece.
Evidence has mostly come from math and coding, where pretraining may already expose models to many solution strategies.
Agents make this a revealing test. They must execute structured actions, follow tool interfaces, learn from feedback, and stay coherent across turns—the very behaviors instruction tuning and RL are widely believed to teach.
Does post-training make successful agentic trajectories newly reachable—or make existing ones more likely?
Post-training raises reliability.
The space of solutions can shrink.
Measure how much post-training
reduces the gains from extra tries.
Explore on hard prompts.
Stay precise on easy ones.
03 / Observation
Large language monkeys.
Now with tools.
Given enough attempts, the base model in the harness often solves more tasks. Post-training pushes per-task success rates towards 100% or 0%.
More tries. More solutions.
Less room for a second chance.
Tasks across 128 attempts
87.6% → 30.0%of tasks are sometimes solved.
Drag the budget to see the crossover. “Never” and “every time” describe the 128 observed rollouts, not absolute impossibility or certainty.
04 / Diagnostic
Put a number on
the lost scalability.
Sharpening Tax compares how much base and post-trained policies gain from retries, across the whole budget.
Gemma 4 · 31B / WebShop
Shaded area = S(K). Both axes are normalized by the available headroom and budget.
TaxS = SBase − SPost
S normalizes the area by budget and first-try failure rate.
model–benchmark pairs have
TaxS(128) > 0.
A positive tax can reflect lost coverage or faster saturation. Read it alongside accuracy and coverage.
The metric, precisely
A(K) = ∑k=1K−1 [pass@K − pass@k]
S(K) = A(K) / [(K − 1)(1 − pass@1)]
Raw mode sums the coverage gaps on a linear attempt axis. Calibrated mode plots recovered headroom, (pass@k − pass@1)/(1 − pass@1), against normalized budget (k − 1)/(K − 1); its area is S(K). S requires pass@1 < 1. The paper’s 36/42 result covers all 14 pairs and three benchmarks.
05 / Solution
One policy.
The right temperature.
Posterior-tempered group sampling (PTGS) heats hard prompts and cools easy ones during RL training.
Explore more paths.
Training onlyMove the slider. PTGS adjusts exploration per prompt.
A wider search, when it helps.
Illustrative policyMore accurate.
Broader coverage. Less tax.
46.5→61.1%
55.0→69.7%
0.094→0.081
Fixed-temperature RL → RL with PTGS. Qwen2.5-7B-Instruct · five-run means · final PPO checkpoints. Same fixed evaluation temperature.
How PTGS chooses the temperature
01 · Draw from the posterior
02 · Turn the draw into a temperature
Each prompt keeps discounted success and failure counts. A draw p̂ from its Beta posterior sets T = τh(p̂), with h(p) = (p̃ − p)/p̃ below the pivot p̃, and (p̃ − p)/(1 − p̃) above it. The example posterior is Beta(2,10); drawing again also updates the policy above.
The demo uses τ = 1.5 and p̃ = 0.5, so T = 1.51−2p̂. Training ramps p̃ from 0.25 to 0.5 with forgetting factor γ = 0.95. These illustrations show the rule, not measured task outcomes. Evaluation uses a fixed temperature.
06 / Theoretical insight
Why the story fits.
Only retries that reach success within K contribute.
A(K) averages these contributions over tasks and rollouts.
How much of success comes from retries?
A(K) counts failed attempts that precede a first success within budget. Later successes contribute more. S(K) normalizes this contribution among first-try failures. A positive tax means successes rely less on retries—through lost coverage or earlier success.
A(K) = E[(H − 1) · 1{H ≤ K}]
Collapse 50% of tasks → lose 50% of raw scalability.
Collapse removes the value of retries.
At p = 1, the first attempt succeeds. At p = 0, no retry can help. If a fraction λ collapses to either extreme independently of base success probability, it removes the same fraction of raw scalability—even when accuracy improves.
TaxA(K) = λ · ABase(K)
Illustration · 4 independent rollouts per group.
Positive: ≥1 success. Mixed: success + failure.
More successes. More useful contrast.
When its tempering raises a hard prompt’s success rate toward ½, PTGS increases the chance of both positive and mixed groups. PPO gets successes to reinforce; GRPO gets reward contrasts that yield nonzero relative advantages.
P(positive) = 1 − (1 − p)n
P(mixed) = 1 − pn − (1 − p)n
07 / Looking ahead
Reliability and reach
should grow together.
On the road to Super Intelligence, the next breakthrough may be the unusual attempt that works.
Measure accuracy. Preserve the possibility.