RL only needs a
domain-specific head.
Fit the current rollout distribution.
An RL run can train the draft head it needs.GrowMTP turns rollout verification into online supervision, growing a draft head from scratch while accelerating the same training run.
Reinforcement learning (RL) post-training drives the frontier capabilities of large language models, with its wall-clock dominated by autoregressive rollout generation. Speculative decoding is an established remedy for this bottleneck, but existing draft heads must be pretrained or warmed up before RL, introducing substantial training cost outside the RL run to be accelerated. We observe that RL training itself provides both conditions required for online draft-head training: its rollout distribution is far narrower than that of pretraining, and its verification step continuously produces supervision signals aligned with this distribution. Building on these observations, we propose GrowMTP, which uses this supervision to train a draft head from scratch entirely within the RL loop, with all head updates detached from the policy backbone. On Qwen3-4B (no draft head), MiMo-7B-SFT (weak head), and Qwen3.5-4B-Base (strong head), GrowMTP achieves rollout speedups of 2.13×, 1.93×, and 1.36×, and end-to-end speedups of 1.60×, 1.41×, and 1.20×, respectively. GrowMTP therefore serves existing RL training frameworks as a modular component, particularly offering a from-scratch acceleration path for models without pretrained draft heads.
A better head accelerates subsequent rollouts
↵Speculative decoding accelerates rollouts, but drafting capability usually has to be trained before RL begins.
| Existing approach | Where the head comes from | Preparation before RL |
|---|---|---|
| Native MTPDeepSeek-V3 |
Joint backbone pretraining |
14.8Tpretraining tokens† |
| EAGLE-style head70B target model |
Dedicated offline head training |
~5,000GPU·h |
Drafting capability enters RL as a prerequisite.
† 14.8T is the full backbone pretraining corpus, not the incremental cost of the MTP head. The two examples use different measures.
A focused rollout distribution and valid verification supervision.
Fit the current rollout distribution.
× The first rejected position still provides a valid signal.
Valid supervision from the first cycle.
Grow the draft head within RL.
RL already provides both the distribution to fit and the supervision to learn from.
RL provides the supervision. Using it for multi-step drafting requires resolving three challenges.
Verification signals belong to specific draft states. Training on the final response can change cycle boundaries, recurrent states, and KV context.
A deeper draft contributes only if all earlier drafts are accepted. Learning at different depths is therefore coupled.
Verification also produces signals after the first rejection, where later drafts depend on a token the target discarded.
p₂ still supervises a valid prefix; p₃ conditions on rejected d₂.
01 State consistency · Rollout vs. teacher forcing
Draft tokens first; train on the committed sequence afterward.
Reconstruct draft states, optimize the acceptance chain, and learn up to the first rejection.
Replay each cycle from its recorded boundary using the original drafts. Recompute hidden states and KV to pair each head prediction with the verification signal under the same context.
Recorded drafts become inputs. No resampling.
Optimize an acceptance-chain surrogate so deeper predictions contribute through earlier acceptance.
αi = 1 − TV(pi, qi) · K = drafting depth
Include the first rejected position and exclude the later suffix. DCA’s soft weights alone do not enforce this rejection boundary.
The first position always teaches the head, even from a random start.
Recorded rollout → draft-path reconstruction
Record each cycle, then reconstruct its draft path.
Backbone features are detached from head training.
A randomly initialized head accelerates the same RL run, with head-update costs included.
Qwen3-4B · Random draft head, K=5 · 500 RL steps · 8 × H800
Single-step time crosses below AR at step 30 on Math and step 22 on Code. These are per-step crossovers.
| Method | K | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | |||
| AR RL | 0 | 1.00 | 523.86 | 1.00× | 723.51 | 1.00× |
| GrowMTP | 5 | 2.91 | 245.95 | 2.13× | 452.53 | 1.60× |
| Method | K | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | |||
| AR RL | 0 | 1.00 | 771.22 | 1.00× | 1028.39 | 1.00× |
| GrowMTP | 5 | 2.64 | 398.65 | 1.93× | 659.39 | 1.56× |
Held-out policy quality remains comparable, while acceptance carries over to new prompts in each training domain.
| Benchmark | AR RL · K=0 | GrowMTP · K=5 | ||
|---|---|---|---|---|
| Mean@16 | τ | Mean@16 | τ | |
| AMC23 | 62.81 | 1.00 | 64.69 | 2.33 |
| AIME24 | 20.42 | 1.00 | 20.42 | 2.34 |
| AIME25 | 18.75 | 1.00 | 21.04 | 2.30 |
| Benchmark | AR RL · K=0 | GrowMTP · K=5 | ||
|---|---|---|---|---|
| Pass@4 | τ | Pass@4 | τ | |
| AtCoder | 42.19 | 1.00 | 43.36 | 2.27 |
| LeetCode | 42.79 | 1.00 | 43.24 | 2.11 |
| LiveCodeBench | 42.37 | 1.00 | 43.51 | 2.22 |
At nearly the same acceptance (2.92 vs. 2.91), GrowMTP uses 9.4% less total compute than the offline-trained route over 500 RL steps.
Detach the head's gradients from the policy, and train it on an objective aligned with acceptance. A rising τ on its own is not evidence of a healthy run.
Qwen3.5-4B-Base · Pretrained head, K=3 · 200 RL steps · Baseline: the same head frozen
Acceptance and policy quality on the same 200-step timeline. Read them together: the arm with the highest τ is the one whose quality goes to zero.
Both detached arms accelerate the run. Among the strategies that keep the policy intact, DCA leads every training-side column.
| Training mode | Loss | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | |||
| Frozen | – | 2.65 | 357.62 | 1.00× | 554.42 | 1.00× |
| Joint collapsed | CE | 3.99 | 372.87 | 0.96× | 643.76 | 0.86× |
| Detached | CE | 3.18 | 277.14 | 1.29× | 507.91 | 1.09× |
| Detached | DCA | 3.22 | 263.86 | 1.36× | 461.56 | 1.20× |
| Training mode | Loss | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | |||
| Frozen | – | 2.59 | 500.09 | 1.00× | 778.54 | 1.00× |
| Joint collapsed | CE | 3.42 | 258.34 | 1.94× | 411.88 | 1.89× |
| Detached | CE | 3.06 | 426.88 | 1.17× | 719.40 | 1.08× |
| Detached | DCA | 3.11 | 408.89 | 1.22× | 694.00 | 1.12× |
Quality separates the strategies. DCA reaches the highest τ on all six benchmarks among the arms that keep their policy; joint CE scores 0.00 on every one.
| Strategy | AMC23 | AIME24 | AIME25 | |||
|---|---|---|---|---|---|---|
| Mean@16 | τ | Mean@16 | τ | Mean@16 | τ | |
| Frozen | 80.47 | 2.61 | 55.00 | 2.57 | 46.67 | 2.57 |
| Joint CE | 0.00 | 3.99 | 0.00 | 3.99 | 0.00 | 3.99 |
| Detached CE | 84.69 | 3.19 | 57.92 | 3.10 | 52.50 | 3.10 |
| Detached DCA | 85.78 | 3.31 | 59.58 | 3.23 | 51.46 | 3.22 |
| Strategy | AtCoder | LeetCode | LiveCodeBench | |||
|---|---|---|---|---|---|---|
| Pass@4 | τ | Pass@4 | τ | Pass@4 | τ | |
| Frozen | 63.29 | 2.52 | 72.52 | 2.52 | 67.39 | 2.52 |
| Joint CE | 0.00 | 3.98 | 0.00 | 3.98 | 0.00 | 3.98 |
| Detached CE | 61.63 | 3.15 | 71.40 | 3.13 | 65.97 | 3.14 |
| Detached DCA | 63.29 | 3.22 | 70.50 | 3.22 | 66.54 | 3.22 |
GrowMTP extends a single-step head to multi-step drafting. Acceptance increases with depth, while end-to-end speedup peaks at K=5 in this setting.
MiMo-7B-SFT · Pretrained single-step head · K=3, 5, 7 · 200 RL steps · Baseline: frozen K=1
Deeper drafting yields higher acceptance throughout training. Step time improves across all trained depths, but its ranking does not follow τ.
K=5 gives the best end-to-end acceleration: 1.41× on Math and 1.35× on Code. K=7 accepts more tokens, but its extra cost reduces the step speedup.
| Method | K | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | |||
| Frozen | 1 | 1.65 | 318.66 | 1.00× | 523.02 | 1.00× |
| GrowMTP | 3 | 3.25 | 176.26 | 1.81× | 378.88 | 1.38× |
| GrowMTP | 5 | 4.04 | 164.84 | 1.93× | 371.02 | 1.41× |
| GrowMTP | 7 | 4.38 | 171.24 | 1.86× | 402.20 | 1.30× |
| Method | K | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | |||
| Frozen | 1 | 1.59 | 441.81 | 1.00× | 707.42 | 1.00× |
| GrowMTP | 3 | 3.11 | 264.69 | 1.67× | 528.64 | 1.34× |
| GrowMTP | 5 | 3.72 | 256.86 | 1.72× | 525.31 | 1.35× |
| GrowMTP | 7 | 4.00 | 252.03 | 1.75× | 542.02 | 1.31× |
The acceptance gains transfer to held-out benchmarks, with comparable policy quality. Every trained depth reaches τ above 3.0 on all six benchmarks.
| Drafting depth | AMC23 | AIME24 | AIME25 | |||
|---|---|---|---|---|---|---|
| Mean@16 | τ | Mean@16 | τ | Mean@16 | τ | |
| Frozen K=1 | 90.47 | 1.67 | 51.88 | 1.64 | 43.13 | 1.64 |
| GrowMTP K=3 | 90.47 | 3.20 | 53.33 | 3.15 | 42.29 | 3.14 |
| GrowMTP K=5 | 91.25 | 3.95 | 56.87 | 3.84 | 46.25 | 3.82 |
| GrowMTP K=7 | 89.69 | 4.30 | 56.04 | 4.11 | 43.33 | 4.08 |
| Drafting depth | AtCoder | LeetCode | LiveCodeBench | |||
|---|---|---|---|---|---|---|
| Pass@4 | τ | Pass@4 | τ | Pass@4 | τ | |
| Frozen K=1 | 69.93 | 1.61 | 80.18 | 1.61 | 74.50 | 1.61 |
| GrowMTP K=3 | 70.10 | 3.02 | 78.38 | 3.03 | 73.84 | 3.03 |
| GrowMTP K=5 | 71.59 | 3.59 | 78.60 | 3.59 | 74.69 | 3.59 |
| GrowMTP K=7 | 69.93 | 3.83 | 76.35 | 3.85 | 72.80 | 3.84 |
Reconstruction lowers training cost, DCA improves multi-step acceptance, and VGM adds further efficiency by enforcing the rejection boundary.
MiMo-7B-SFT · Math reasoning · Pretrained single-step head · 200 RL steps · K=5, with K=1 also evaluated for DCA
Similar aggregate acceptance, lower training cost. Reconstruction reaches τ=4.04 versus 4.02 for teacher forcing, while reducing end-to-end step time from 408.09s to 371.02s.
| Configuration | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | ||
| Teacher forcing | 4.02 | 169.05 | 1.89× | 408.09 | 1.28× |
| GrowMTP | 4.04 | 164.84 | 1.93× | 371.02 | 1.41× |
| Configuration | AMC23 | AIME24 | AIME25 | |||
|---|---|---|---|---|---|---|
| Mean@16 | τ | Mean@16 | τ | Mean@16 | τ | |
| Teacher forcing | 91.56 | 3.93 | 54.37 | 3.82 | 43.96 | 3.79 |
| GrowMTP | 91.25 | 3.95 | 56.87 | 3.84 | 46.25 | 3.82 |
The acceptance advantage emerges at multi-step depth. At K=1 the four objectives yield similar τ; at K=5, DCA leads in training acceptance, end-to-end speedup, and all three held-out acceptance columns.
| Configuration | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | ||
| CE · K=1 | 1.84 | 274.47 | 1.16× | 507.99 | 1.03× |
| KL · K=1 | 1.81 | 288.58 | 1.10× | 532.91 | 0.98× |
| TV · K=1 | 1.83 | 288.30 | 1.11× | 530.29 | 0.99× |
| DCA · K=1 | 1.82 | 280.58 | 1.14× | 517.00 | 1.01× |
| CE · K=5 | 3.86 | 175.70 | 1.81× | 384.22 | 1.36× |
| KL · K=5 | 3.85 | 174.98 | 1.82× | 382.41 | 1.37× |
| TV · K=5 | 3.94 | 174.00 | 1.83× | 382.69 | 1.37× |
| DCA · K=5 | 4.04 | 164.84 | 1.93× | 371.02 | 1.41× |
| Configuration | AMC23 | AIME24 | AIME25 | |||
|---|---|---|---|---|---|---|
| Mean@16 | τ | Mean@16 | τ | Mean@16 | τ | |
| CE · K=1 | 90.16 | 1.88 | 51.67 | 1.87 | 40.42 | 1.87 |
| KL · K=1 | 90.31 | 1.86 | 54.37 | 1.85 | 44.58 | 1.84 |
| TV · K=1 | 89.22 | 1.88 | 56.04 | 1.86 | 43.13 | 1.86 |
| DCA · K=1 | 89.53 | 1.87 | 53.96 | 1.85 | 36.88 | 1.86 |
| CE · K=5 | 89.84 | 3.81 | 54.17 | 3.67 | 43.54 | 3.66 |
| KL · K=5 | 89.69 | 3.67 | 53.75 | 3.53 | 42.29 | 3.51 |
| TV · K=5 | 89.53 | 3.80 | 55.00 | 3.67 | 40.83 | 3.66 |
| DCA · K=5 | 91.25 | 3.95 | 56.87 | 3.84 | 46.25 | 3.82 |
Masking post-rejection positions improves end-to-end efficiency under both objectives. With VGM, step speedup rises from 1.28× to 1.37× under KL, and from 1.35× to 1.41× under DCA.
| Configuration | τ | Rollout | End-to-end step | ||
|---|---|---|---|---|---|
| Time (s) | Speedup | Time (s) | Speedup | ||
| KL · VGM off | 3.80 | 186.40 | 1.71× | 408.65 | 1.28× |
| KL · VGM on | 3.85 | 174.98 | 1.82× | 382.41 | 1.37× |
| DCA · VGM off | 4.00 | 165.97 | 1.92× | 387.27 | 1.35× |
| DCA · VGM on | 4.04 | 164.84 | 1.93× | 371.02 | 1.41× |
| Configuration | AMC23 | AIME24 | AIME25 | |||
|---|---|---|---|---|---|---|
| Mean@16 | τ | Mean@16 | τ | Mean@16 | τ | |
| KL · VGM off | 88.59 | 3.66 | 54.17 | 3.49 | 41.04 | 3.48 |
| KL · VGM on | 89.69 | 3.67 | 53.75 | 3.53 | 42.29 | 3.51 |
| DCA · VGM off | 89.22 | 3.94 | 52.92 | 3.81 | 41.88 | 3.79 |
| DCA · VGM on | 91.25 | 3.95 | 56.87 | 3.84 | 46.25 | 3.82 |
Three model starting points, the potential of larger RL runs, and the domain specialization of an online-trained head.
GrowMTP adapts to different starting points, from a randomly initialized head to pretrained draft heads. The three evaluated models all gain end-to-end acceleration on mathematical reasoning.
Random initialization
1.60×
End-to-end step speedup
K=5 · 500 RL steps
Baseline
Autoregressive decoding
Pretrained single-step head
1.41×
End-to-end step speedup
K=5 · 200 RL steps
Baseline
Frozen pretrained head, K=1
Pretrained head
1.20×
End-to-end step speedup
K=3 · 200 RL steps
Baseline
Frozen pretrained head, K=3
Longer rollouts and longer runs could increase end-to-end speedup by increasing rollout’s share of runtime and amortizing early head-training overhead.
Projection from one measured run
The Qwen3-4B Math run measures 1.60× at 500 steps and an 8K rollout length. Under the paper’s assumptions, the projection reaches 1.91× at 1,000 steps / 16K and 2.17× at 2,000 steps / 32K.
Qwen3-4B · DAPO-Math-17K · K=5 · AR baseline
| RL steps | 8K rollout | 16K rollout | 32K rollout |
|---|---|---|---|
| 500 | 1.60×Measured | 1.79×Projected | 1.93×Projected |
| 1,000 | 1.70×Projected | 1.91×Projected | 2.08×Projected |
| 2,000 | 1.76×Projected | 1.99×Projected | 2.17×Projected |
The grown head specializes to the rollout distribution it learns from. Acceptance falls on held-out data and drops further across domains.
Qwen3-4B · K=5 · Two heads grown from scratch for 500 RL steps, then frozen for evaluation
| Head training domain | Training rollout | Same-domain evaluation | Cross-domain evaluation |
|---|---|---|---|
| Math | 2.91 | 2.32 | 1.45 |
| Code | 2.64 | 2.20 | 1.75 |
| Draft head | AMC23 | AIME24 | AIME25 | Avg. |
|---|---|---|---|---|
| Math-trained | 2.33 | 2.34 | 2.30 | 2.32 |
| Code-trained | 1.80 | 1.79 | 1.68 | 1.75 |
| Draft head | AtCoder | LeetCode | LiveCodeBench | Avg. |
|---|---|---|---|---|
| Math-trained | 1.50 | 1.40 | 1.46 | 1.45 |
| Code-trained | 2.27 | 2.11 | 2.22 | 2.20 |
A concentrated rollout distribution lets online training grow a useful head for the current RL run. General serving must cover unforeseen deployment distributions; the paper therefore retains a role for pretrained MTP heads. The cross-domain acceptance drop makes this specialization visible.
A randomly initialized draft head learns entirely within RL, without offline pretraining or warm-up.
On Qwen3-4B, end-to-end training speeds up by 1.60× on Math and 1.56× on Code, including head updates.
Online adaptation improves both random and pretrained heads, with comparable held-out quality in the evaluated runs.