Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MotionVLA: Injecting Geometric Motion into
Vision-Language-Action Model

Shanglin Yuan1,2 · Weiheng Zhao1,2 · Xianda Guo2,3 · Wei Sui2,† · Li Yu1 · Wenyu Liu1 · Xinggang Wang1,‡

1Huazhong University of Science and Technology · 2D-Robotics · 3Wuhan University

†Project Lead · ‡Corresponding Author

Accepted to Conference on Robot Learning (CoRL), 2026

Paper on arXiv Project Page MotionVLA models on Hugging Face Apache-2.0 License

MotionVLA equips pi0-family policies with a strictly past-only RGB motion history. A frozen TraceAnything encoder supplies compact trajectory-field tokens; current visual tokens retrieve task-relevant motion through Decouple, while historical end-effector reconstruction grounds the representation in control-relevant dynamics. Recouple provides an optional path for injecting the retrieved motion back into the vision-language-action stream.

Contents

Installation

Clone the repository and initialize TraceAnything:

git clone --recurse-submodules https://github.com/hustvl/MotionVLA.git
cd MotionVLA
git submodule update --init --recursive third_party/TraceAnything

The pi0 and pi0.5 integrations both provide the openpi Python package and must be installed in separate Python 3.11 environments. Select one family:

# pi0
conda create -n motionvla-pi0 python=3.11 -y
conda activate motionvla-pi0
python -m pip install -e pi0/openpi
# pi0.5
conda create -n motionvla-pi05 python=3.11 -y
conda activate motionvla-pi05
python -m pip install -e pi05/openpi

Model Preparation

Set the paths used by the pi0 RoboTwin example:

export MOTIONVLA_ROOT="$(pwd)"
export TRACE_CHECKPOINT=/absolute/path/to/trace_anything.pt
export DATASET_ROOT=/absolute/path/to/stack_blocks_three
export NORM_STATS_DIR=/absolute/path/to/stack_blocks_three_stats
export RUN_DIR=/absolute/path/to/motionvla_stack_blocks_three
export ROBOTWIN_ROOT=/absolute/path/to/RoboTwin
Component Source Local target
TraceAnything source ByteDance-Seed/TraceAnything third_party/TraceAnything/
TraceAnything checkpoint trace_anything.pt $TRACE_CHECKPOINT
RoboTwin RoboTwin-Platform/RoboTwin $ROBOTWIN_ROOT

The TraceAnything checkpoint must match the SHA-256 recorded in the selected YAML config. Task-specific modules use random initialization with seed 42.

Data Preparation

The example below uses the single RoboTwin task stack_blocks_three. Prepare a 10 Hz LeRobot dataset with the following fields:

Field Shape / type Description
observation.images.cam_high RGB video Head-camera observation
observation.images.cam_left_wrist RGB video Left wrist camera
observation.images.cam_right_wrist RGB video Right wrist camera
observation.state [14] Dual-arm joints and grippers
action [14] Dual-arm joint targets and grippers
observation.eef_pose [14] Left/right xyz + wxyz in the SAPIEN world frame
episode_index, frame_index integer Episode-local temporal indices
task string Language instruction

norm_stats.json must contain statistics for state, actions, and eef_pose. Action statistics are computed after converting the twelve joint channels to state-relative deltas; gripper channels remain absolute.

Training

The following command trains stack_blocks_three with online TraceAnything feature extraction. Run it from the repository root in the pi0 environment:

export XLA_PYTHON_CLIENT_PREALLOCATE=false

python -m pi0.scripts.train_robotwin \
  --config pi0/configs/robotwin_stack_blocks_three_finetune.yaml \
  --dataset-root "$DATASET_ROOT" \
  --norm-stats-dir "$NORM_STATS_DIR" \
  --traceanything-root "$MOTIONVLA_ROOT/third_party/TraceAnything" \
  --trace-checkpoint "$TRACE_CHECKPOINT" \
  --checkpoint-dir "$RUN_DIR" \
  --task-name stack_blocks_three

Evaluation

Start the policy server in the pi0 environment:

export XLA_PYTHON_CLIENT_PREALLOCATE=false

python -m pi0.scripts.serve_robotwin \
  --run-dir "$RUN_DIR" \
  --checkpoint-step 30000 \
  --traceanything-root "$MOTIONVLA_ROOT/third_party/TraceAnything" \
  --trace-checkpoint "$TRACE_CHECKPOINT" \
  --port 8000

In the RoboTwin environment, run the benchmark evaluator with the matching family adapter:

cd "$ROBOTWIN_ROOT"

PYTHONPATH="$MOTIONVLA_ROOT:$MOTIONVLA_ROOT/pi0/openpi/src" \
python script/eval_policy.py \
  --config "$MOTIONVLA_ROOT/pi0/configs/robotwin_stack_blocks_three_eval.yaml"

The adapter maintains the strictly past-only 10 Hz history and sends each 14D absolute joint target to RoboTwin's existing Base_Task.take_action() path.

Additional Benchmarks

The pi0 and pi0.5 READMEs contain the LIBERO training, TraceAnything-cache, serving, and evaluation commands. The custom ordered-contact RoboTwin task is documented in robotwin_touch.

Acknowledgments

We thank the RoboTwin authors for providing the simulation benchmark used in our experiments.

License

MotionVLA-authored code is licensed under Apache-2.0. Third-party components retain their respective terms; see third_party/THIRD_PARTY_NOTICES.md.

Citation

If you find MotionVLA useful in your work, please consider citing our paper:

@inproceedings{yuan2026motionvla,
  title     = {MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model},
  author    = {Yuan, Shanglin and Zhao, Weiheng and Guo, Xianda and Sui, Wei and Yu, Li and Liu, Wenyu and Wang, Xinggang},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026}
}

About

[CoRL2026] MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

Resources

Stars

18 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages