InternRobotics/F1-VLA

Model

controls autoplay muted playsinline loop width="720">

32

4 commits

5 linked in READMEs

updated Sep 9, 2025

See the code

README

🏁 Best viewed with sound on

F1: A Vision Language Action Model Bridging
Understanding and Generation to Actions

Paper Code Website

🚀 Key Innovations

  • 🧠 Predictive Inverse Dynamics: Visual foresight generation for planning-based control
  • 🏗️ Mixture-of-Transformer: Three specialized experts (Understanding, Generation, Action)
  • 📈 Three-Stage Training: Progressive alignment, pretraining, and adaptation

🤖 Real-World Robot Experiments

Diverse manipulation tasks across multiple robot platforms.

📊 Performance Summary

TaskPlatformF1π0Improvement
Multi-taskGenie-182.2%65.2%+17.0%
AdaptationFranka66.7%53.3%+13.4%
Long-horizonARX LIFT II40.0%0.0%+40.0%
Dynamic EnvARX LIFT II66.7%33.3%+33.4%

Usage

Please refer to our official repo F1-VLA.

📚 Citation

If you find our work helpful, please cite:

@article{f1_vla_2025,
  title={F1: A Vision Language Action Model Bridging Understanding and Generation to Actions},
  author={Qi Lv and Weijie Kong and Hao Li and Jia Zeng and Zherui Qiu and Delin Qu and Haoming Song and Qizhi Chen and Xiang Deng and Jiangmiao Pang},
  journal={Conference/Journal Name},
  year={2025},
  url={https://arxiv.org/abs/2509.06951}
}

License

This work is under the cc-by-nc-sa-4.0.

Acknowledgements

This repository is based on Lerobot, Any4lerobot, and VAR.

custom_code
feature-extraction
manipulation
robotics
safetensors
transformers
vision-language-model

Contributors

ZE
Zeng-Jia

3 commits

Jia-Zeng

1 commits

InternRobotics/F1-VLA

Model

controls autoplay muted playsinline loop width="720">

32

4 commits

5 linked in READMEs

updated Sep 9, 2025

See the code

README

🏁 Best viewed with sound on

F1: A Vision Language Action Model Bridging
Understanding and Generation to Actions

Paper Code Website

🚀 Key Innovations

  • 🧠 Predictive Inverse Dynamics: Visual foresight generation for planning-based control
  • 🏗️ Mixture-of-Transformer: Three specialized experts (Understanding, Generation, Action)
  • 📈 Three-Stage Training: Progressive alignment, pretraining, and adaptation

🤖 Real-World Robot Experiments

Diverse manipulation tasks across multiple robot platforms.

📊 Performance Summary

TaskPlatformF1π0Improvement
Multi-taskGenie-182.2%65.2%+17.0%
AdaptationFranka66.7%53.3%+13.4%
Long-horizonARX LIFT II40.0%0.0%+40.0%
Dynamic EnvARX LIFT II66.7%33.3%+33.4%

Usage

Please refer to our official repo F1-VLA.

📚 Citation

If you find our work helpful, please cite:

@article{f1_vla_2025,
  title={F1: A Vision Language Action Model Bridging Understanding and Generation to Actions},
  author={Qi Lv and Weijie Kong and Hao Li and Jia Zeng and Zherui Qiu and Delin Qu and Haoming Song and Qizhi Chen and Xiang Deng and Jiangmiao Pang},
  journal={Conference/Journal Name},
  year={2025},
  url={https://arxiv.org/abs/2509.06951}
}

License

This work is under the cc-by-nc-sa-4.0.

Acknowledgements

This repository is based on Lerobot, Any4lerobot, and VAR.

custom_code
feature-extraction
manipulation
robotics
safetensors
transformers
vision-language-model

Contributors

ZE
Zeng-Jia

3 commits

Jia-Zeng

1 commits