R1 MobileVLA-R1 2.0

Vision · Language · Action

MobileVLA-R1 2.0:
RL-Enhanced Reasoning for Mobile Robot Control

Ting Huang1*, Yue Huang2*, Zeyu Zhang1*†, Shuicheng Yan3, Hao Tang1‡

1School of Computer Science, Peking University   2SCUT   3National University of Singapore

*Equal contribution. Project lead. Corresponding author.

00 /

Method

Reasoning to action
Multi-granularity chain-of-thought data engine diagram
Chapter 01

Multi-granularity CoT data engine.

Given multimodal observations, language instructions, and optional state-action histories, the engine generates episode-, step-, and navigation-level reasoning traces together with executable targets, followed by automatic parsing and semi-automatic quality verification.

MobileVLA-R1 2.0 structured reasoning and action decoder diagram
Chapter 02

MobileVLA-R1 2.0

Multimodal observations and language instructions are fused for structured reasoning, which conditions a task-level action decoder to predict locomotion (Vx, Vy, ω) and a discrete task-level behavior primitive α for mobile robot control.

GRPO-based reasoning-to-action optimization diagram
Chapter 03

GRPO-based reasoning-to-action optimization.

For each multimodal input, the policy samples multiple structured outputs whose induced actions are evaluated by movement, behavior, and format rewards. Group-relative advantages, together with KL regularization to a frozen reference policy, are then used for offline policy optimization.

01 /

Real World Eval

Ego + Exo synchronized views