DexPolicy

Policy optimization for dexterous control

DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

Haoyu Wang1*   Siyuan Qian1*   Yanjun Li1*   Zeyu Zhang1*†   Yandong Guo2   Boxin Shi1   Hao Tang1‡
1School of Computer Science, Peking University    2AI2 Robotics
*Equal contribution.  †Project lead.  ‡Corresponding author.

Method

Scheduled exploration
DexPolicy method overview

DexPolicy Overview

(A) Simulation and real-robot tasks. (B) Gaussian noise scheduled by training steps; the PPO example does not represent within-trial phases or other methods' schedules. (C) Paired updates and deterministic target success (%) from the baseline to DexPolicy. Simulation averages five shared objects over three seeds; hardware averages three objects with 20 trials from one model per condition. Policies are object-specific, with distinct simulation and hardware embodiments; GRPO uses continuation budgets.

Learning Curves

Learning histories on the three hardware object geometries, using separate benchmark policies. PPO and FPO show stochastic evaluation returns; GRPO shows continuation rollout returns. Shading indicates sample standard deviation over three seeds; GRPO includes development seed 0. Dotted lines mark GRPO phase reloads; PPO checkpoints are nominal. Compare conditions within panels: return measurements and budgets differ across methods.

DexPolicy learning curves on three hardware object geometries

MuJoCo Simulation

Eight evaluation rollouts
Rollout 01
Rollout 02
Rollout 03
Rollout 04
Rollout 05
Rollout 06
Rollout 07
Rollout 08

Real World Evaluation

Three real-world demonstrations
Demonstration 01
Demonstration 02
Demonstration 03