FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation

Abstract

Real-world reinforcement learning for robotic manipulation remains challenging, and this difficulty is amplified for flow matching policies: applying policy gradient methods to these policies is fundamentally limited by the need to backpropagate through time (BPTT) along the multi-step ODE that maps noise to actions, which is computationally prohibitive and numerically fragile. We propose FlowDPG, a DDPG-style method specifically designed for flow matching policies that distills the critic gradient into the velocity field at training time, bypassing BPTT entirely. Intuitively, FlowDPG combines two complementary vectors: the demonstration-driven velocity that keeps the action feasible, and the critic-driven correction that steers it toward higher value. Our contributions are threefold: (1) a BPTT-free in-place framework that leaves the deployed multi-step inference path unchanged, (2) a formal connection between the FlowDPG update direction and vanilla Deterministic Policy Gradient via three explicit approximations, and (3) real-world validation on two long-horizon, multi-stage, dual-arm tasks on two different robots — AirPods assembly on a Franka and scrambled-egg cooking on a YAM.

Method Overview

Classical flow matching transports noise $\epsilon$ to a demonstration-feasible action $a$ via $v_\theta$, and lacks the ability to explore new actions outside the demonstration distribution. FlowDPG adds a critic-driven correction along $\nabla_a Q(s, a)$ to reach a value-improved target $a^*$, then distills the resulting velocity $v'_\theta$ back into the flow field; in this process, new actions with high value may be explored.

Scrambled Eggs on Yam arms (New)

14 stages · tool use · irreversible spills · exogenous dynamics · force-limited failure

Airpods Assembly on Franka arms

One hour uncut · 25 AirPods cases assembled back to back · no human intervention

8 stages · bimanual coordination · millimetre insertion tolerance · randomised poses · no stage resets · compounding failure

Recovery from Unseen Failures

Stage-Aware Reward Model

Successful rollout

Failed rollout

SARM progress overlaid on each episode. Left: it advances through the stages. Right: it stalls, then drops at the failure — and that is what the critic is updated on.

Results

Rubric score

AirPods (Franka) 0 20 40 60 80 100 Overall success (%) 62 BC 66 68 FQL 78 80 QAM 86 90 FlowDPG AirPods (Franka) 0 20 40 60 80 100 Rubric score (%) 66.58 BC 68.08 70.17 FQL 79.42 81.42 QAM 88.17 92.25 FlowDPG Scrambled eggs (YAM) 0 20 40 60 80 100 Rubric score (%) 64.24 BC 68.86 FQL 76.90 QAM 85.76 FlowDPG
offlineafter online fine-tuningerror bars: s.e. over 5 checkpoints

Overall: episodes completing every stage. Rubric: 0–3 per stage as a percentage of the available points (24 AirPods, 42 egg; criteria below). Protocol: 50 rollouts per cell, 10 at each of 5 checkpoints; ± is the standard error over the five checkpoint means. Egg task is offline-only.

Full baseline sweep (Airpods)

Comparison against prior RL methods 0 20 40 60 80 100 End-to-end success (%) BC (base) 64 Value-conditioning RA-BC 76 AWR 76 RECAP 72 Auxiliary module DSRL 68 PLD 76 RLT 80 Critic gradient QAM 80 FlowDPG (offline) 88 FlowDPG (+online) 92

25 cases (50 pod manipulations) per method, reporting the best end-to-end success rate.

Ablations (Airpods)

(a) Projection Error 0 0.02 0.04 0.06 0.08 0.1 0 2 4 6 8 10 Projection error Training steps (k) w/o consistency w/ consistency (b) Q-value Discrepancy 0 0.1 0.2 0.3 0.4 0.5 0 2 4 6 8 10 Q discrepancy Training steps (k) w/o consistency w/ consistency (c) Training Loss 0 0.04 0.08 0.12 0.16 0 2 4 6 8 10 Training loss Training steps (k) w/o adaptive shift w/ adaptive shift (d) Success Rate 0 20 40 60 80 100 Success rate (%) 72 w/o both 76 w/o adaptive shift 80 w/o consistency 92 Full FlowDPG

Without the consistency term the projection error stays an order of magnitude higher and the critic is queried out of distribution (92 → 80). Without the adaptive shift the update is unbounded and the loss oscillates (92 → 76). Both address the same cause, so removing them together adds little (72). Lines are a moving average, shading ±1 SD.

Stage coverage under reward ablation 100% 96% 96% 96% 96% 92% 92% 92% Full reward 100% 96% 92% 88% 84% 84% 80% 80% w/o progress reward 100% 96% 92% 88% 84% 80% 76% 76% w/o stage-transition reward 100% 80% 76% 72% 68% 64% 60% 60% Only terminal reward Grasp case Open case Grasp R pod Insert R pod Grasp L pod Insert L pod Close case Place case 60 80 100

Each variant fails differently. Without the progress term the value flattens within a stage. Without the stage-transition bonus the policy lingers instead of closing a sub-task out. With only the terminal bonus it skips the pod-handling sequence entirely.

Evaluation Rubrics

0–3 per stage; a stage never reached scores 0. Criteria fixed before scoring.

AirPods assembly (Franka), 8 stages, 24 points

Stage 3 – Successful 2 – Minor error 1 – Major error 0 – Failed
Grasp caseSecure grasp, correct orientation, no slipGrasp with adjustment or delayCase lifted but shifted or dropped onceFails to grasp case
Open caseLid fully opened in one clean motionOpens with adjustment or hesitationOpens only after repeated attempts, or case displacedLid not opened
Grasp right podSecure grasp, correct approach directionGrasp with adjustment or delayWrong approach direction, pod dislodged then re-graspedFails to grasp pod
Insert right podSeated on the first attempt, fully inSeated after one retrySeated after two or more retries, or insertion dislodges the other podPod not seated
Grasp left podSecure grasp, correct approach directionGrasp with adjustment or delayWrong approach direction, pod dislodged then re-graspedFails to grasp pod
Insert left podSeated on the first attempt, fully inSeated after one retrySeated after two or more retries, or insertion dislodges the right podPod not seated
Close caseLid closed cleanly with both pods seatedCloses after minor adjustmentCloses with a pod unseated, or needs repositioningCase not closed
Place casePlaced upright within the target regionPlaced with minor misalignmentDropped or placed outside the regionCase not placed

Scrambled eggs (YAM), 14 stages, 42 points

Pick up and crack occur twice, once per egg: 14 scored stages from 12 types.

Stage 3 – Successful 2 – Minor error 1 – Major error 0 – Failed
Pick up eggSecure grasp, no slip or damageGrasp with adjustment or delayLifts egg but drops or damages itFails to grasp egg
Crack eggClean crack, proper shell separationMinor shell fragments or awkward motionIncorrect or messy crackFails to crack egg
Put saltCorrect amount applied accuratelyMinor spillage or hesitationIncorrect amount or placementSalt not applied, or bottle dropped
Pick up forkSecure and correct graspGrasp with adjustmentUnstable or incorrect graspFails to pick up fork, or drops it
Whisk eggThorough whisking with correct motionPartial or inefficient whiskingMinimal or incorrect motionNo whisking
Place fork backPlaced neatly in correct locationMinor misalignmentDropped or poorly placedFork not placed back
Put eggs into panClean transfer into panMinor spillage or hesitationSignificant spillageEggs not transferred
Pick up spatulaSecure and correct graspGrasp with adjustmentUnstable or incorrect graspFails to pick up spatula
Stir eggsProper stirring across panIncomplete or uneven stirringMinimal or incorrect stirringNo stirring
Transfer eggsClean and accurate transferMinor spillage or inefficiencyPartial or messy transferEggs not transferred
Place spatula backPlaced correctly and safelyMinor misplacementDropped or unsafe placementSpatula not placed back
Serve eggsServed correctly and fullyMinor error or delayIncorrect or incomplete servingEggs not served