Under review · ICLR 2027

TD-MTP: Coupling Structured Planning with Multi-Step Value Learning

We study how the planner in a model-based agent interacts with the value target its critic learns from. Across ten paired tasks, the effect of switching to a tensor-graph planner depends on whether the critic regresses onto a one-step or a TD(λ) target.

The paper is under double-blind review. The manuscript and anonymized code will be linked here once they are available.

8 of 10
tasks where the target alone beats the baseline
1 of 10
tasks where the planner alone beats it
2 of 10
interactions clear of zero
10 / 22 / 6
wins / ties / losses on 38 tasks

components alone

Effect of each component on its own

Each row is a task, and each mark is the change in final return from adding one component to the baseline, as a share of that task’s baseline return so that tasks with different reward scales share one axis. The multi-step target finishes above the baseline on eight of the ten tasks and the structured planner on one.

-50% -25% +25% +50% +75% +100% +125% 0 baseline humanoid-run humanoid-walk acrobot-swingup pole run maze walk stand hurdle stair
planner alone multi-step target alone
Five seeds per cell. The humanoid-run baseline is low (72.9), so its target effect of +97 return appears as +133% and runs off the scale. The target finishes below the baseline on stair and hurdle, and stair is also the task where the two components together beat both.

h1-stair

Each component alone and both together on h1-stair

On h1-stair, the tensor-graph planner and the multi-step target each finish below the baseline when added alone, and the agent with both finishes well above it.

Final return, mean over the last three evaluations, averaged over five seeds. Both planners in a block share their model initialization.

all ten tasks

Stair is the only task with this pattern

Each panel is one paired task, with each arm plotted against that task’s baseline. Only on stair do both components finish below the baseline alone and above it together. The combination finishes ahead of the baseline on seven of the ten tasks and behind it on hurdle, run and walk.

humanoid-run P T P+T humanoid-walk P T P+T acrobot-swingup P T P+T stair P T P+T pole P T P+T stand P T P+T hurdle P T P+T run P T P+T walk P T P+T maze P T P+T P = planner alone · T = target alone · P+T = both · dashed line = that task’s baseline · teal above, red below
The vertical scale is set per panel, so directions can be compared across panels but magnitudes only within one. Marks below the dashed line finish below that task’s baseline.

planner × target

The planner’s effect under each target

Each of the ten task blocks was trained four times, with the planner off or on and a one-step or TD(λ) target. Each mark is the change in final return from turning the planner on, under the selected target.

interaction per task

Interaction estimates with 95% intervals

The interaction is the planner’s effect under the multi-step target minus its effect under the one-step target, with a 95% interval over the five registered seeds. Eight of the ten estimates are positive. The interval excludes zero on stair and pole, and on the other eight tasks the data cannot distinguish the interaction from zero.

-200 +200 +400 +600 +800 0 interaction in return units (scales differ between tasks) stair 5/5 pole 5/5 hurdle 3/5 humanoid-walk 4/5 humanoid-run 3/5 acrobot-swingup 3/5 stand 4/5 maze 3/5 run 2/5 walk 2/5
interval excludes zero interval contains zero right column: seeds out of five with a positive interaction
The intervals are unadjusted paired Student-t intervals on a contrast chosen after seeing the data, so they describe spread and should not be read as a hypothesis test. Reward scales differ by an order of magnitude between tasks, so the magnitudes are not averaged.

proposed mechanism

A possible mechanism, which we have not tested

The planner is also the policy that collects data, so a proposal that reaches a different solution puts a different kind of trajectory into the replay buffer. A one-step target uses only the first reward of that trajectory and bootstraps the rest from the critic’s current estimate, so the later rewards the planner found reach the critic only indirectly.

A multi-step target over the stored transitions puts those rewards into the target directly. This is a hypothesis. The paper’s own controls attribute more of the effect to the elite fit than to the proposals, and the isolated contribution of tensor candidates is −14.6 [−107, +78], which is consistent with zero. Testing the hypothesis would need critic error measured on selected and unselected candidates, which we have not run.

locomotion and manipulation

The target alone on locomotion and manipulation tasks

These counts score the multi-step target on its own, on every task where a target-only arm was trained. It wins four and loses two of the ten locomotion tasks, and wins three and loses six of the sixteen manipulation tasks.

paired study · 10 return tasks 4 4 2 manipulation suites · 16 success tasks 3 7 6 wins · ties · losses, under the paper’s own tie tolerances
A tie means the two arms differ by less than 3% of baseline return or 0.05 success. The band is descriptive and is not an equivalence test. Two of the four benchmark suites have no target-only arm and are not counted here.

across four benchmarks

Results on 38 tasks

10
wins
22
ties
6
losses
38
tasks

The tasks come from HumanoidBench, DMControl, Meta-World and MyoSuite. The largest gains are on tasks where the baseline is weak, such as balance_hard (175 → 430), stair (490 → 664) and KeyTurn, which the baseline never solves and TD-MTP solves at 0.84. Most ties are on tasks where the baseline already saturates.

Most losses are on manipulation tasks, where the target helps some tasks a lot and hurts others. On locomotion the combination seems to improve reliability more than speed. On h1-pole, TD-MTP crosses the 890-return threshold on all seven seeds and the target alone on five, and both have the same median crossing time of 0.76M steps.

rollouts

Grasping rollouts on a dexterous hand

These clips replay one TD-MTP grasping policy, trained for 4.89M environment steps, in a separate simulator on twelve objects. In the lift scene the hand picks the object off a table and carries it to a goal above it. In the two-tray scene it picks the object from one tray and drops it into a second tray 0.22 m away.

Success rates over 64 episodes per object, from one block of seeds (3000–3063). Lift counts episodes where the object stays at least 0.15 m above the table for 10 consecutive steps. Place counts episodes that end with the object at rest near the goal and clear of its support, and drop counts episodes that end with it at rest inside the destination tray. Rates below 0.5 are in red. The policy reads the true object pose instead of an estimate, so a deployed controller would score lower. Each clip is the best episode from a separate 8-episode run, and its label says whether that run succeeded at least once.
The two-tray result depends on how far apart the trays are. With the trays 0.32 m apart the commanded palm position is outside the arm’s reach, the arm lurches and often flings the object, and the mean drop rate is 0.16. At 0.22 m every object improves, though cube_medium and sphere_large still mostly fail, and in the sphere_large clip the hand keeps its grip past the release point.

training budget

At 4M steps TD-MTP falls below the baseline on stair

On stair, doubling the training budget to 4M steps moves TD-MTP from above the baseline to below it. The interaction stays positive at both budgets, so the planner still helps more under the multi-step target than under the one-step target, but at 4M the combined agent returns less than the baseline.

2M · five seeds 294 +planner 460 +target 651 TD-MTP baseline 525 interaction +423 TD-MTP +126 · above baseline 4M · three seeds 495 +planner 777 +target 705 TD-MTP baseline 755 interaction +187 TD-MTP -50 · below baseline
The two budgets use different run sets and the 4M comparison has three seeds, so this result cannot separate a budget effect from seed variability or from the change of run set. We include it because it matters to anyone reusing the method.

practical notes

Notes for practitioners

01

Check which value target the critic uses before evaluating a multi-modal proposal. Under a one-step target a richer proposal can lower return, and an ablation will then count against the planner when the cause is the pairing.

02

On a sequential buffer, handle termination and truncation separately. Treating a time limit as a terminal state drops the value of a real successor state, with a bias that grows with how often sampled windows cross an episode boundary, and the training loss shows no sign of it.

03

Tune λ per task. On reach, return falls from 9053 to 7422 to 5240 as λ goes from 0 to 0.5 to 0.8. Longer returns help when bootstrap error is larger than the mismatch between the policy that collected the data and the prior the target bootstraps with.

04

Use λ = 0 as the baseline arm. At λ = 0 the gated return reduces term by term to the standard one-step target, so the baseline arm is the reference agent itself and does not need a separate reimplementation.

The full method, proofs and per-task results are in the manuscript.

Every number on this page outside the hand rollouts is recomputed from the manuscript’s tables and was checked against the 26 September draft. The rollout rates come from a separate sim-to-sim evaluation run on 12 September and are not in the manuscript.