Highlight

Method

Full-step and naive few-step actions compared with RFPO's rectified flow policy for reliable few-step execution.
Few-step execution of reward-optimized flow policies. Naive step reduction can create a discretization gap between full-step and few-step actions. RFPO rectifies student-induced transport paths and uses complementary action-space supervision to enable reliable few-step execution.
RFPO on-policy training with an online flow student, frozen PPO teacher, reward-aware online Reflow, and action distillation, followed by one-step deployment.
Overview of RFPO. During on-policy training, student-induced full-step endpoints are rectified by reward-aware online Reflow, while multi-budget student actions are distilled toward a frozen PPO controller. An adaptive budget objective further regularizes the fidelity–compute trade-off. At deployment, all auxiliary components are removed and only the flow student is retained for one-step Euler execution.
Normalized RFPO action-state densities over flow time for hip and thigh joints across all four legs of Unitree Go2.
Action-flow geometry of RFPO on Unitree Go2. Normalized action-state densities over flow time for representative hip and thigh joints across all four legs. The learned policy exhibits concentrated and smoothly evolving transport patterns across joint dimensions.

Sim2Real on Unitree G1

Simulation
Real World
Simulation
Real World

Sim2Real on Unitree Go2 (1-Step)

Simulation
Real World
Simulation
Real World
Simulation
Real World

Sim2Real on Unitree Go2 (64-Step)

Simulation
Real World
Simulation
Real World
Simulation
Real World

BibTeX

@misc{huang2026rfpo,
  title  = {RFPO: Rectified Flow Policy Optimization for Embodied Control},
  author = {Ting Huang and Lisiyu Pan and Haoyu Wang and Zeyu Zhang and Siyuan Qian and Yanjun Li and Yandong Guo and Boxin Shi and Hao Tang},
  year   = {2026}
}