FYX026/MergeVLA-LIBERO

Model

Model Card for MergeVLA-LIBERO

2

4 commits

1 linked in READMEs

updated Dec 8, 2025

See the code

README

Model Card for MergeVLA-LIBERO

MergeVLA — Single-Skill Experts for Spatial / Object / Goal / Long-10 (LIBERO Task Suite). These models are used as the base expert checkpoints for our MergeVLA.

Model Details

Each uploaded model is a 0.68B-parameter VLA model (excluding the vision backbone) composed of:

  • Qwen2.5-0.5B as the Vision-Language Model (VLM)
  • A lightweight 0.18B Action Expert
  • A two-layer Proprioceptive Projector MLP

✔️ Performance (Success Rates on LIBERO)

Task FamilySuccess Rate (%)
Spatial98.0
Object98.6
Goal95.0
Long-1095.0

🧠 Training Details

Each expert is fine-tuned independently using modified LIBER demonstrations in RLDS format.

CategoryValue
LoRAEnabled (rank = 64)
OptimizerAdamW
Learning Rate2e-4
Batch Size8 (×2 grad accumulation)
num_images_in_input2

Training Steps

  • Spatial — 30,000
  • Object — 20,000
  • Goal — 30,000
  • Long-10 — 50,000

Citation instructions

@misc{fu2025mergevla,
      title={MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent}, 
      author={Yuxia Fu and Zhizhen Zhang and Yuqi Zhang and Zijian Wang and Zi Huang and Yadan Luo},
      year={2025},
      eprint={2511.18810},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2511.18810}, 
}
safetensors

FYX026/MergeVLA-LIBERO

Model

Model Card for MergeVLA-LIBERO

2

4 commits

1 linked in READMEs

updated Dec 8, 2025

See the code

README

Model Card for MergeVLA-LIBERO

MergeVLA — Single-Skill Experts for Spatial / Object / Goal / Long-10 (LIBERO Task Suite). These models are used as the base expert checkpoints for our MergeVLA.

Model Details

Each uploaded model is a 0.68B-parameter VLA model (excluding the vision backbone) composed of:

  • Qwen2.5-0.5B as the Vision-Language Model (VLM)
  • A lightweight 0.18B Action Expert
  • A two-layer Proprioceptive Projector MLP

✔️ Performance (Success Rates on LIBERO)

Task FamilySuccess Rate (%)
Spatial98.0
Object98.6
Goal95.0
Long-1095.0

🧠 Training Details

Each expert is fine-tuned independently using modified LIBER demonstrations in RLDS format.

CategoryValue
LoRAEnabled (rank = 64)
OptimizerAdamW
Learning Rate2e-4
Batch Size8 (×2 grad accumulation)
num_images_in_input2

Training Steps

  • Spatial — 30,000
  • Object — 20,000
  • Goal — 30,000
  • Long-10 — 50,000

Citation instructions

@misc{fu2025mergevla,
      title={MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent}, 
      author={Yuxia Fu and Zhizhen Zhang and Yuqi Zhang and Zijian Wang and Zi Huang and Yadan Luo},
      year={2025},
      eprint={2511.18810},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2511.18810}, 
}
safetensors