Under review · ICLR 2027
We study how the planner in a model-based agent interacts with the value target its critic learns from. Across ten paired tasks, the effect of switching to a tensor-graph planner depends on whether the critic regresses onto a one-step or a TD(λ) target.
The paper is under double-blind review. The manuscript and anonymized code will be linked here once they are available.
components alone
Each row is a task, and each mark is the change in final return from adding one component to the baseline, as a share of that task’s baseline return so that tasks with different reward scales share one axis. The multi-step target finishes above the baseline on eight of the ten tasks and the structured planner on one.
h1-stair
On h1-stair, the tensor-graph planner and the multi-step target each finish below the baseline when added alone, and the agent with both finishes well above it.
all ten tasks
Each panel is one paired task, with each arm plotted against that task’s baseline. Only on stair do both components finish below the baseline alone and above it together. The combination finishes ahead of the baseline on seven of the ten tasks and behind it on hurdle, run and walk.
planner × target
Each of the ten task blocks was trained four times, with the planner off or on and a one-step or TD(λ) target. Each mark is the change in final return from turning the planner on, under the selected target.
interaction per task
The interaction is the planner’s effect under the multi-step target minus its effect under the one-step target, with a 95% interval over the five registered seeds. Eight of the ten estimates are positive. The interval excludes zero on stair and pole, and on the other eight tasks the data cannot distinguish the interaction from zero.
proposed mechanism
The planner is also the policy that collects data, so a proposal that reaches a different solution puts a different kind of trajectory into the replay buffer. A one-step target uses only the first reward of that trajectory and bootstraps the rest from the critic’s current estimate, so the later rewards the planner found reach the critic only indirectly.
A multi-step target over the stored transitions puts those rewards into the target directly. This is a hypothesis. The paper’s own controls attribute more of the effect to the elite fit than to the proposals, and the isolated contribution of tensor candidates is −14.6 [−107, +78], which is consistent with zero. Testing the hypothesis would need critic error measured on selected and unselected candidates, which we have not run.
A structured proposal can only reach regions its control points cover. The bound on this is tight at the setting used in the paper.
Proposition 2 →How a multi-step target on a sequential buffer has to treat termination and truncation, and the bias it picks up, with no sign in the loss, when a time limit is treated as terminal.
Proposition 3 →locomotion and manipulation
These counts score the multi-step target on its own, on every task where a target-only arm was trained. It wins four and loses two of the ten locomotion tasks, and wins three and loses six of the sixteen manipulation tasks.
across four benchmarks
The tasks come from HumanoidBench, DMControl, Meta-World and MyoSuite. The largest gains are on tasks where the baseline is weak, such as balance_hard (175 → 430), stair (490 → 664) and KeyTurn, which the baseline never solves and TD-MTP solves at 0.84. Most ties are on tasks where the baseline already saturates.
Most losses are on manipulation tasks, where the target helps some tasks a lot and hurts others. On locomotion the combination seems to improve reliability more than speed. On h1-pole, TD-MTP crosses the 890-return threshold on all seven seeds and the target alone on five, and both have the same median crossing time of 0.76M steps.
rollouts
These clips replay one TD-MTP grasping policy, trained for 4.89M environment steps, in a separate simulator on twelve objects. In the lift scene the hand picks the object off a table and carries it to a goal above it. In the two-tray scene it picks the object from one tray and drops it into a second tray 0.22 m away.
training budget
On stair, doubling the training budget to 4M steps moves TD-MTP from above the baseline to below it. The interaction stays positive at both budgets, so the planner still helps more under the multi-step target than under the one-step target, but at 4M the combined agent returns less than the baseline.
practical notes
Check which value target the critic uses before evaluating a multi-modal proposal. Under a one-step target a richer proposal can lower return, and an ablation will then count against the planner when the cause is the pairing.
On a sequential buffer, handle termination and truncation separately. Treating a time limit as a terminal state drops the value of a real successor state, with a bias that grows with how often sampled windows cross an episode boundary, and the training loss shows no sign of it.
Tune λ per task. On reach, return falls from 9053 to 7422 to 5240 as λ goes from 0 to 0.5 to 0.8. Longer returns help when bootstrap error is larger than the mismatch between the policy that collected the data and the prior the target bootstraps with.
Use λ = 0 as the baseline arm. At λ = 0 the gated return reduces term by term to the standard one-step target, so the baseline arm is the reference agent itself and does not need a separate reimplementation.
The full method, proofs and per-task results are in the manuscript.
Every number on this page outside the hand rollouts is recomputed from the manuscript’s tables and was checked against the 26 September draft. The rollout rates come from a separate sim-to-sim evaluation run on 12 September and are not in the manuscript.