On the possibilities we keep—and the ones we lose

Sharpening Tax
in Post-Training

A better first try. A smaller world of solutions?

* Work done at Meta

01 / Background & motivation

New capabilities.
Or sharper probabilities?

The sharpening hypothesis: post-training amplifies behaviors already present in the base model, improving first-try accuracy while narrowing solution coverage.

Distribution sharpening Schematic: post-training concentrates probability on a dominant behavior already present in the base model, while suppressing other existing modes. Both distributions are normalized; no new mode is added. BasePost-trainedSchematic Possible behaviorsProbability

The missing piece.

Evidence has mostly come from math and coding, where pretraining may already expose models to many solution strategies.

Agents make this a revealing test. They must execute structured actions, follow tool interfaces, learn from feedback, and stay coherent across turns—the very behaviors instruction tuning and RL are widely believed to teach.

Does post-training make successful agentic trajectories newly reachable—or make existing ones more likely?

02 / The project in three stepsObserve

Post-training raises reliability.
The space of solutions can shrink.

Diagnose

Measure how much post-training
reduces the gains from extra tries.

Improve

Explore on hard prompts.
Stay precise on easy ones.

03 / Observation

Large language monkeys.
Now with tools.

A large language monkey, in the harness A minimal monkey at a keyboard uses a tool terminal: the base model working through a harness. tools

Given enough attempts, the base model in the harness often solves more tasks. Post-training pushes per-task success rates towards 100% or 0%.

More tries. More solutions.

BasePost-trained
128

Less room for a second chance.

Tasks across 128 attempts

Never solvedSometimesEvery time

87.6% → 30.0%of tasks are sometimes solved.

Measured · 500 tasks · 128 rollouts per task

Drag the budget to see the crossover. “Never” and “every time” describe the 128 observed rollouts, not absolute impossibility or certainty.

04 / Diagnostic

Put a number on
the lost scalability.

Sharpening Tax compares how much base and post-trained policies gain from retries, across the whole budget.

How much do retries recover?

Gemma 4 · 31B / WebShop

Base
Post-trained
128

Shaded area = S(K). Both axes are normalized by the available headroom and budget.

TaxS = SBase − SPost

—−—=—

S normalizes the area by budget and first-try failure rate.

36/42

model–benchmark pairs have
TaxS(128) > 0.

A positive tax can reflect lost coverage or faster saturation. Read it alongside accuracy and coverage.

The metric, precisely

A(K) = ∑k=1K−1 [pass@K − pass@k]
S(K) = A(K) / [(K − 1)(1 − pass@1)]

Raw mode sums the coverage gaps on a linear attempt axis. Calibrated mode plots recovered headroom, (pass@k − pass@1)/(1 − pass@1), against normalized budget (k − 1)/(K − 1); its area is S(K). S requires pass@1 < 1. The paper’s 36/42 result covers all 14 pairs and three benchmarks.

05 / Solution

One policy.
The right temperature.

Posterior-tempered group sampling (PTGS) heats hard prompts and cools easy ones during RL training.

Track successes & failuresSample difficultySet temperatureRoll out & update

Explore more paths.

Training only
1.38temperature
10%
Hard promptEasy prompt

Move the slider. PTGS adjusts exploration per prompt.

A wider search, when it helps.

Illustrative policy
Fixed T = 1PTGS

More accurate.
Broader coverage. Less tax.

Accuracy pass@1 ↑

46.561.1%

Coverage pass@128 ↑

55.069.7%

TaxS(128) ↓

0.0940.081

Fixed-temperature RL → RL with PTGS. Qwen2.5-7B-Instruct · five-run means · final PPO checkpoints. Same fixed evaluation temperature.

How PTGS chooses the temperature

01 · Draw from the posterior

Example: 1 success + 9 failures, with a Beta(1,1) prior.

02 · Turn the draw into a temperature

Low p̂ → heat up. High p̂ → cool down.

Each prompt keeps discounted success and failure counts. A draw p̂ from its Beta posterior sets T = τh(p̂), with h(p) = (p̃ − p)/p̃ below the pivot p̃, and (p̃ − p)/(1 − p̃) above it. The example posterior is Beta(2,10); drawing again also updates the policy above.

The demo uses τ = 1.5 and p̃ = 0.5, so T = 1.51−2p̂. Training ramps p̃ from 0.25 to 0.5 with forgetting factor γ = 0.95. These illustrations show the rule, not measured task outcomes. Evaluation uses a fixed temperature.

06 / Theoretical insight

Why the story fits.

Outcome · K = 4Contribution
First try
✓···
0
Retry succeeds
×××✓
3
No success
××××
0

Only retries that reach success within K contribute.
A(K) averages these contributions over tasks and rollouts.

How much of success comes from retries?

A(K) counts failed attempts that precede a first success within budget. Later successes contribute more. S(K) normalizes this contribution among first-try failures. A positive tax means successes rely less on retries—through lost coverage or earlier success.

A(K) = E[(H − 1) · 1{H ≤ K}]

07 / Looking ahead

Reliability and reach
should grow together.

On the road to Super Intelligence, the next breakthrough may be the unusual attempt that works.
Measure accuracy. Preserve the possibility.

Explore the code