RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts

Hongbin Lin1Chaoda Zheng2Yiming Yang1Xiangyu Li2Shijia Chen2Jinhao Deng2Kangjie Chen2
Dongbin Zhang2Jie Feng3Yu Zhang2Xianming Liu2Shuguang Cui1Boyang Wang2Zhen Li1
1 The Chinese University of Hong Kong Shenzhen (cuhk.edu.cn)2 XPeng Motors (xpeng.com)3 Xi'an University of Electronic Science and Technology (xidian.edu.cn)

RoXDrive Demo

For efficient web delivery, these videos use the compressed versions.

Qualitative Visualization

Public Benchmark

nuScenes

Public nuScenes · Original vs RoXDrive.

Scaling Study

In-House Driving Data

Closed-loop comparisons across challenging in-house driving scenes.

Collision Avoidance

Safe responses to dynamic obstacles and collision risks.

Keep in Drivable Area

Lane-boundary awareness and recovery within the drivable region.

TL;DR

We introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization through action-faithful world-model rollouts. RoXDrive consists of two stages:

  1. 1

    Model pre-training

    Alongside imitation pre-training, the Action-Vision Faithfulness Evaluator uses geometry-aware trajectory supervision to verify long-horizon action-vision consistency.

  2. 2

    Action-faithful RL post-training

    With a fixed world model, only action-faithful long-horizon rollouts receive dense safety-aware scoring for scene-level closed-loop RL.

Comparison of conventional simulator-based reinforcement learning, vanilla multi-step world-model rollouts, and RoXDrive action-faithful closed-loop reinforcement learning
RoXDrive selects realistic, long-horizon, action-faithful rollouts for dense safety-aware scoring and scene-level policy optimization.

Motivation

Video world models are promising scalable simulators for closed-loop policy learning, but visual realism alone does not guarantee that future observations causally follow the policy's ego actions.

Fine-tuning only marginally reduces relative-motion errors; small inverse-dynamics inaccuracies still accumulate into jitter and long-horizon drift, making rollouts unreliable for policy optimization.

Motivation examples showing accumulated jitter errors and long-horizon drift in fine-tuned video world-model rollouts

Method

RoXDrive model pre-training and action-faithful reinforcement learning post-training pipeline

Illustration of RoXDrive. 1) Model pre-training (top): Beyond policy pre-training, we train the action-vision faithfulness evaluator (AVFE) with our geometry-aware auxiliary trajectory supervision. 2) Action-faithful RL post-training (bottom): Given the initial scene, agents iteratively interact with a frozen world model to form long-horizon scene rollouts, retaining only action-faithful ones via AVFE for dense safety-aware scoring and scene-level closed-loop RL post-training.

Quantitative Results

Open-loop collision rates (2 s) and closed-loop evaluation on 4,675 nuScenes clips (4 s, two action steps).

MethodRLCol. 1s ↓Col. 2s ↓Col. Avg. ↓Obj. Col. ↓Lane Viol. ↓Progress ↑Comfort ↑Clearance ↑Driving Score ↑
ST-P3×0.230.620.432685570.3970.4250.3870.341
VAD×0.070.170.128208240.2550.6810.4180.271
UniAD×0.620.580.603488700.4080.8050.4150.430
CLEAR‡✓0.110.230.17——————
Drive-r1‡✓0.020.060.04——————
DiffusionDrive×0.0680.0730.0702667720.6940.8660.3920.526
DiffusionDrive + RoXDrive✓0.0290.0630.0461895620.7170.8830.3800.588
SparseDrive×0.0000.0440.0221976230.7160.8910.3840.580
SparseDrive + RoXDrive✓0.0000.0150.0071735980.7750.9220.3880.622

Citation

@misc{lin2026roxdrive,
  title     = {RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts},
  author    = {Hongbin Lin and Chaoda Zheng and Yiming Yang and Xiangyu Li and Shijia Chen and Jinhao Deng and Kangjie Chen and Dongbin Zhang and Jie Feng and Yu Zhang and Xianming Liu and Shuguang Cui and Boyang Wang and Zhen Li},
  year      = {2026}
}