ABot-PhysWorld is a physically consistent, action-controllable video world model for robotic manipulation, built on a 14-billion-parameter Diffusion Transformer. It integrates physics-aware training, memory-efficient preference optimization, and precise spatial action injection to generate realistic and physically plausible robot-object interactions โ even in zero-shot settings.
training/README_A2V.md and inference/README_A2V.md.training/README_DPO.md.training/.EZS-Bench/.inference/.Industrial-Grade Data Pipeline
Curated ~3M real-world manipulation clips from five datasets (AgiBot, RoboCoin, RoboMind, Galaxea, OXE) with motion, semantic, and action consistency filtering, plus hierarchical sampling for balanced generalization.
Physics-Aware DPO Training
Introduces a decoupled VLM-based discriminator: Qwen3-VL generates task-specific physics checklists, Gemini 3 Pro scores videos via Chain-of-Thought; combined with LoRA-augmented DPO on a 14B DiT to enforce physical plausibility.
Parallel Context Blocks for Action Control
Enables precise action-conditioned generation by residually injecting spatial action maps into cloned DiT blocks, preserving physical priors while supporting cross-embodiment control.
EZSbench โ First True Zero-Shot Benchmark
Fully training-independent evaluation covering unseen robot, scene, and task combinations, with dual-model scoring to eliminate self-evaluation bias.
Embodied-ZeroShot Benchmark for Physically Consistent Video Generation ๐คโจ
EZS-Bench is a zero-shot evaluation benchmark designed to rigorously assess physically plausible video generation in robotic manipulation. It evaluates models on physical consistency, action controllability, and cross-embodiment generalizationโwith no training-test data overlap. ๐๐ฌ
โ
True Zero-Shot Evaluation
Unseen combinations of:
๐จ Dual-Source Data Construction
๐ง Physics-Aware Evaluation
๐ Comprehensive Metrics
Evaluates:
Download evaluation data from ModelScope:
git lfs install
git clone https://www.modelscope.cn/datasets/amap_cvlab/EZS-Bench_data.git
Install and run the evaluation toolkit:
cd EZS-Bench
pip install -e .
# Full evaluation (Video Quality + Domain Score)
torchrun --standalone --nproc_per_node=4 evaluate_ezsbench.py \
--data_file /path/to/EZS-Bench_data/video_prompt_question_196_ezs0.jsonl \
--method_name "YourMethod" \
--method_dir /path/to/generated_videos \
--output_dir ./results
The VLM judge model (Qwen2.5-VL-72B-Instruct, ~150 GB) is automatically downloaded on first run.
๐ See EZS-Bench/README.md for full documentation.
We evaluate ABot-PhysWorld on three key aspects:
| Capability | Benchmark | Ours | Best Baseline | Gain |
|---|---|---|---|---|
| Physical Fidelity | PBench (Domain Score) | 0.9306 | 0.8644 (Wan2.5) | +6.62% |
| Zero-Shot Generalization | EZSbench (Domain Score) | 0.8366 | 0.7951 (WoW) | +4.15% |
| Action Control | Trajectory Consistency | 0.8522 | 0.8157 (Enerverse) | +3.65% |
โ ABot-PhysWorld establishes a new standard for physically grounded, controllable, and generalizable world models in robotic manipulation.
Selected representative zero-shot generation results demonstrating ABot-PhysWorld's strong generalization and physical plausibility.
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
Note: The Gemini watermark (bottom-right) indicates the initial frame generated by Gemini (ensuring it is completely unseen); all other frames are generated by ABot-PhysWorld.
![]() | ![]() |
![]() | ![]() |
Note: The Gemini watermark (bottom-right) indicates the initial frame generated by Gemini (ensuring it is completely unseen); all other frames are generated by ABot-PhysWorld.
![]() | ![]() |
![]() | ![]() |
We conducted systematic qualitative comparative experiments on the PAI-Bench benchmark dataset. Below are the generated results from several typical scenarios.
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
| Task | Baselines | Ours |
|---|---|---|
| Grasping | Frequent penetration, floatation | โ Firm contact, no violation |
| Long-horizon Planning | Inconsistent state transitions | โ Coherent multi-step reasoning |
| Rigid-body Dynamics | Unphysical deformations | โ Preserved geometry and mass behavior |
| Contact Modeling | Non-contact attraction | โ Realistic interaction onset |
Our model consistently generates physically valid trajectories even in complex, unseen scenarios โ proving its utility as a reliable simulator for embodied AI.
Generate physically plausible robot manipulation videos using the ABot-PhysWorld fine-tuned model.
# Create conda environment
conda create -n abot-physworld python=3.10
conda activate abot-physworld
# Install PyTorch with CUDA support
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
# Install dependencies
pip install -r requirements.txt
Hardware Requirements:
| Configuration | VRAM | Notes |
|---|---|---|
| Recommended | >= 60GB | Best performance, no tiling needed |
| Minimum | >= 24GB | Uses tiled VAE (enabled by default) |
cd inference
# Download demo data and run inference
python inference.py \
--jsonl_path assets/demo.jsonl \
--output_dir ./outputs/demo \
--save_first_frames
This generates videos for 2 Franka robot manipulation samples. The model checkpoint is auto-downloaded from ModelScope on first run.
python inference.py \
--input_image /path/to/image.jpg \
--prompt "robot arm picks up the red cube from the table" \
--output_dir ./outputs
Prepare a JSONL file (each line is a sample):
{"video": "path/to/image.jpg", "prompt": "robot grasps the object"}
{"video": "path/to/image2.jpg", "prompt": "robot places object on table"}
Then run:
python inference.py \
--jsonl_path data.jsonl \
--output_dir ./outputs \
--num_samples 100 # Process max 100 samples
python inference.py --help
Key parameters:
--checkpoint_path: Local path to model weights (auto-downloads if not provided)--cache_dir: Directory to store downloaded weights (default: ./checkpoints)--height, --width: Video resolution (default: 480ร832)--num_frames: Number of output frames (default: 81 โ 5.4s at 15fps)--num_inference_steps: Denoising steps, higher = better quality but slower (default: 50)--cfg_scale: Classifier-free guidance scale (default: 5.0)--seed: Random seed for reproducibility--gpu_id: GPU device index{output_dir}/{image_name}_generated.mp4{output_dir}/{unique_id}_generated.mp4 + results.json (with status for each sample)Auto-Download: The fine-tuned checkpoint is automatically downloaded from ModelScope on first inference run.
Manual Download (Optional):
pip install modelscope
modelscope download --model amap_cvlab/Abot-PhysWorld --local_dir ./inference/checkpoints
Base Model: Wan2.1-I2V-14B-480P is also auto-downloaded by DiffSynth-Studio.
For detailed setup instructions, examples, and troubleshooting, see inference/README.md.
We provide full-parameter SFT training scripts to fine-tune Wan2.1-I2V-14B-480P on your own robot manipulation datasets.
The v1 SFT training dataset is available on ModelScope:
git lfs install
git clone https://www.modelscope.cn/datasets/amap_cvlab/ABot-PhysWorld_SFT_Training_Data_v1.git
cd training
# Prepare your dataset (JSONL format, see training/assets/demo_train.jsonl)
# Then launch 8-GPU training:
bash run_train.sh
RESUME_CHECKPOINT=./outputs/sft_training/step-800.safetensors \
bash run_train_resume.sh
# First run: train + save encoded features
ENCODED_CACHE_DIR=./encoded_cache bash run_train.sh
# Subsequent runs: reuse cached features (much faster)
ENCODED_CACHE_DIR=./encoded_cache bash run_train.sh
For detailed training instructions, data preparation, and parameter reference, see training/README.md.
We release the A2V training and inference code for action-conditioned video generation via VACE parallel context blocks. Given an input image and an action trajectory (end-effector poses), the model generates a physically consistent video of the robot executing the specified actions.
cd training
# Train VACE module on top of SFT DiT
DIT_CHECKPOINT=/path/to/dit_checkpoint.safetensors \
DATASET_BASE_PATH=/path/to/dataset \
DATASET_METADATA_PATH=/path/to/metadata.jsonl \
bash run_train_a2v.sh
cd inference
# Run A2V inference (checkpoints auto-downloaded from ModelScope)
python inference_a2v.py \
--jsonl_path ./assets/demo_a2v.jsonl \
--output_dir ./outputs/a2v_results
# With trajectory overlay visualization
python inference_a2v.py \
--jsonl_path data.jsonl \
--output_dir ./outputs \
--overlay_action_condition
For detailed documentation, see training/README_A2V.md and inference/README_A2V.md.
We release the DPO (Direct Preference Optimization) training pipeline for physics-aware alignment. Using winner/loser video pairs, the model learns to generate videos that better respect physical laws via LoRA fine-tuning.
cd training
# Step 1: Preprocess DPO data
DPO_JSONL=/path/to/dpo_pairs.jsonl \
CACHE_DIR=/path/to/dpo_cache \
bash run_preprocess_dpo.sh
# Step 2: Train DPO LoRA
DIT_CHECKPOINT=/path/to/dit_checkpoint.safetensors \
DPO_CACHE_DIR=/path/to/dpo_cache \
bash run_train_dpo.sh
For detailed documentation, see training/README_DPO.md.
If you find ABot-PhysWorld is useful in your research or applications, please consider giving us a star ๐ and citing it by the following BibTeX entry:
@article{chen2026abotphysworld,
title={ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment},
author={Yuzhi Chen, Ronghan Chen, Dongjie Huo, Yandan Yang, Dekang Qi, Haoyun Liu, Tong Lin, Shuang Zeng, Junjin Xiao, Xinyuan Chang, Feng Xiong, Xing Wei, Zhiheng Ma, Mu Xu},
journal={arXiv preprint arXiv:2603.23376},
year={2026}
}
This project builds upon the following open-source projects. We thank these teams for their contributions:
956 followers ยท starred Apr 2026
Python
99.0%
ABot-PhysWorld is a physically consistent, action-controllable video world model for robotic manipulation, built on a 14-billion-parameter Diffusion Transformer. It integrates physics-aware training, memory-efficient preference optimization, and precise spatial action injection to generate realistic and physically plausible robot-object interactions โ even in zero-shot settings.
training/README_A2V.md and inference/README_A2V.md.training/README_DPO.md.training/.EZS-Bench/.inference/.Industrial-Grade Data Pipeline
Curated ~3M real-world manipulation clips from five datasets (AgiBot, RoboCoin, RoboMind, Galaxea, OXE) with motion, semantic, and action consistency filtering, plus hierarchical sampling for balanced generalization.
Physics-Aware DPO Training
Introduces a decoupled VLM-based discriminator: Qwen3-VL generates task-specific physics checklists, Gemini 3 Pro scores videos via Chain-of-Thought; combined with LoRA-augmented DPO on a 14B DiT to enforce physical plausibility.
Parallel Context Blocks for Action Control
Enables precise action-conditioned generation by residually injecting spatial action maps into cloned DiT blocks, preserving physical priors while supporting cross-embodiment control.
EZSbench โ First True Zero-Shot Benchmark
Fully training-independent evaluation covering unseen robot, scene, and task combinations, with dual-model scoring to eliminate self-evaluation bias.
Embodied-ZeroShot Benchmark for Physically Consistent Video Generation ๐คโจ
EZS-Bench is a zero-shot evaluation benchmark designed to rigorously assess physically plausible video generation in robotic manipulation. It evaluates models on physical consistency, action controllability, and cross-embodiment generalizationโwith no training-test data overlap. ๐๐ฌ
โ
True Zero-Shot Evaluation
Unseen combinations of:
๐จ Dual-Source Data Construction
๐ง Physics-Aware Evaluation
๐ Comprehensive Metrics
Evaluates:
Download evaluation data from ModelScope:
git lfs install
git clone https://www.modelscope.cn/datasets/amap_cvlab/EZS-Bench_data.git
Install and run the evaluation toolkit:
cd EZS-Bench
pip install -e .
# Full evaluation (Video Quality + Domain Score)
torchrun --standalone --nproc_per_node=4 evaluate_ezsbench.py \
--data_file /path/to/EZS-Bench_data/video_prompt_question_196_ezs0.jsonl \
--method_name "YourMethod" \
--method_dir /path/to/generated_videos \
--output_dir ./results
The VLM judge model (Qwen2.5-VL-72B-Instruct, ~150 GB) is automatically downloaded on first run.
๐ See EZS-Bench/README.md for full documentation.
We evaluate ABot-PhysWorld on three key aspects:
| Capability | Benchmark | Ours | Best Baseline | Gain |
|---|---|---|---|---|
| Physical Fidelity | PBench (Domain Score) | 0.9306 | 0.8644 (Wan2.5) | +6.62% |
| Zero-Shot Generalization | EZSbench (Domain Score) | 0.8366 | 0.7951 (WoW) | +4.15% |
| Action Control | Trajectory Consistency | 0.8522 | 0.8157 (Enerverse) | +3.65% |
โ ABot-PhysWorld establishes a new standard for physically grounded, controllable, and generalizable world models in robotic manipulation.
Selected representative zero-shot generation results demonstrating ABot-PhysWorld's strong generalization and physical plausibility.
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
Note: The Gemini watermark (bottom-right) indicates the initial frame generated by Gemini (ensuring it is completely unseen); all other frames are generated by ABot-PhysWorld.
![]() | ![]() |
![]() | ![]() |
Note: The Gemini watermark (bottom-right) indicates the initial frame generated by Gemini (ensuring it is completely unseen); all other frames are generated by ABot-PhysWorld.
![]() | ![]() |
![]() | ![]() |
We conducted systematic qualitative comparative experiments on the PAI-Bench benchmark dataset. Below are the generated results from several typical scenarios.
![]() | ![]() |
![]() | ![]() |
![]() | ![]() |
| Task | Baselines | Ours |
|---|---|---|
| Grasping | Frequent penetration, floatation | โ Firm contact, no violation |
| Long-horizon Planning | Inconsistent state transitions | โ Coherent multi-step reasoning |
| Rigid-body Dynamics | Unphysical deformations | โ Preserved geometry and mass behavior |
| Contact Modeling | Non-contact attraction | โ Realistic interaction onset |
Our model consistently generates physically valid trajectories even in complex, unseen scenarios โ proving its utility as a reliable simulator for embodied AI.
Generate physically plausible robot manipulation videos using the ABot-PhysWorld fine-tuned model.
# Create conda environment
conda create -n abot-physworld python=3.10
conda activate abot-physworld
# Install PyTorch with CUDA support
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
# Install dependencies
pip install -r requirements.txt
Hardware Requirements:
| Configuration | VRAM | Notes |
|---|---|---|
| Recommended | >= 60GB | Best performance, no tiling needed |
| Minimum | >= 24GB | Uses tiled VAE (enabled by default) |
cd inference
# Download demo data and run inference
python inference.py \
--jsonl_path assets/demo.jsonl \
--output_dir ./outputs/demo \
--save_first_frames
This generates videos for 2 Franka robot manipulation samples. The model checkpoint is auto-downloaded from ModelScope on first run.
python inference.py \
--input_image /path/to/image.jpg \
--prompt "robot arm picks up the red cube from the table" \
--output_dir ./outputs
Prepare a JSONL file (each line is a sample):
{"video": "path/to/image.jpg", "prompt": "robot grasps the object"}
{"video": "path/to/image2.jpg", "prompt": "robot places object on table"}
Then run:
python inference.py \
--jsonl_path data.jsonl \
--output_dir ./outputs \
--num_samples 100 # Process max 100 samples
python inference.py --help
Key parameters:
--checkpoint_path: Local path to model weights (auto-downloads if not provided)--cache_dir: Directory to store downloaded weights (default: ./checkpoints)--height, --width: Video resolution (default: 480ร832)--num_frames: Number of output frames (default: 81 โ 5.4s at 15fps)--num_inference_steps: Denoising steps, higher = better quality but slower (default: 50)--cfg_scale: Classifier-free guidance scale (default: 5.0)--seed: Random seed for reproducibility--gpu_id: GPU device index{output_dir}/{image_name}_generated.mp4{output_dir}/{unique_id}_generated.mp4 + results.json (with status for each sample)Auto-Download: The fine-tuned checkpoint is automatically downloaded from ModelScope on first inference run.
Manual Download (Optional):
pip install modelscope
modelscope download --model amap_cvlab/Abot-PhysWorld --local_dir ./inference/checkpoints
Base Model: Wan2.1-I2V-14B-480P is also auto-downloaded by DiffSynth-Studio.
For detailed setup instructions, examples, and troubleshooting, see inference/README.md.
We provide full-parameter SFT training scripts to fine-tune Wan2.1-I2V-14B-480P on your own robot manipulation datasets.
The v1 SFT training dataset is available on ModelScope:
git lfs install
git clone https://www.modelscope.cn/datasets/amap_cvlab/ABot-PhysWorld_SFT_Training_Data_v1.git
cd training
# Prepare your dataset (JSONL format, see training/assets/demo_train.jsonl)
# Then launch 8-GPU training:
bash run_train.sh
RESUME_CHECKPOINT=./outputs/sft_training/step-800.safetensors \
bash run_train_resume.sh
# First run: train + save encoded features
ENCODED_CACHE_DIR=./encoded_cache bash run_train.sh
# Subsequent runs: reuse cached features (much faster)
ENCODED_CACHE_DIR=./encoded_cache bash run_train.sh
For detailed training instructions, data preparation, and parameter reference, see training/README.md.
We release the A2V training and inference code for action-conditioned video generation via VACE parallel context blocks. Given an input image and an action trajectory (end-effector poses), the model generates a physically consistent video of the robot executing the specified actions.
cd training
# Train VACE module on top of SFT DiT
DIT_CHECKPOINT=/path/to/dit_checkpoint.safetensors \
DATASET_BASE_PATH=/path/to/dataset \
DATASET_METADATA_PATH=/path/to/metadata.jsonl \
bash run_train_a2v.sh
cd inference
# Run A2V inference (checkpoints auto-downloaded from ModelScope)
python inference_a2v.py \
--jsonl_path ./assets/demo_a2v.jsonl \
--output_dir ./outputs/a2v_results
# With trajectory overlay visualization
python inference_a2v.py \
--jsonl_path data.jsonl \
--output_dir ./outputs \
--overlay_action_condition
For detailed documentation, see training/README_A2V.md and inference/README_A2V.md.
We release the DPO (Direct Preference Optimization) training pipeline for physics-aware alignment. Using winner/loser video pairs, the model learns to generate videos that better respect physical laws via LoRA fine-tuning.
cd training
# Step 1: Preprocess DPO data
DPO_JSONL=/path/to/dpo_pairs.jsonl \
CACHE_DIR=/path/to/dpo_cache \
bash run_preprocess_dpo.sh
# Step 2: Train DPO LoRA
DIT_CHECKPOINT=/path/to/dit_checkpoint.safetensors \
DPO_CACHE_DIR=/path/to/dpo_cache \
bash run_train_dpo.sh
For detailed documentation, see training/README_DPO.md.
If you find ABot-PhysWorld is useful in your research or applications, please consider giving us a star ๐ and citing it by the following BibTeX entry:
@article{chen2026abotphysworld,
title={ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment},
author={Yuzhi Chen, Ronghan Chen, Dongjie Huo, Yandan Yang, Dekang Qi, Haoyun Liu, Tong Lin, Shuang Zeng, Junjin Xiao, Xinyuan Chang, Feng Xiong, Xing Wei, Zhiheng Ma, Mu Xu},
journal={arXiv preprint arXiv:2603.23376},
year={2026}
}
This project builds upon the following open-source projects. We thank these teams for their contributions:
956 followers ยท starred Apr 2026
Python
99.0%