variante/llava-1.5-7b-llara-D-inBC-Aux-B-VIMA-80k

Model

LLaRA Model Card

2

5 commits

1 linked in READMEs

updated Jul 15, 2024

See the code

README


LLaRA Model Card

This model is released with paper LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

Xiang Li1, Cristina Mata1, Jongwoo Park1, Kumara Kahatapitiya1, Yoo Sung Jang1, Jinghuan Shang1, Kanchana Ranasinghe1, Ryan Burgert1, Mu Cai2, Yong Jae Lee2, and Michael S. Ryoo1

1Stony Brook University 2University of Wisconsin-Madison

Model details

Model type: LLaRA is an open-source visuomotor policy trained by fine-tuning LLaVA-7b-v1.5 on instruction-following data D-inBC and 4 auxiliary datasets, converted from VIMA-Data. For the conversion code, please refer to convert_vima.ipynb

Model date: llava-1.5-7b-llara-D-inBC-Aux-B-VIMA-80k was trained in June 2024.

Paper or resources for more information: https://github.com/LostXine/LLaRA

Where to send questions or comments about the model: https://github.com/LostXine/LLaRA/issues

Intended use

Primary intended uses: The primary use of LLaRA is research on large multimodal models for robotics.

Primary intended users: The primary intended users of the model are researchers and hobbyists in robotics, computer vision, natural language processing, machine learning, and artificial intelligence.

image-text-to-text
llara
llava
robotics
safetensors
text-generation
transformers

Contributors

variante

5 commits

variante/llava-1.5-7b-llara-D-inBC-Aux-B-VIMA-80k

Model

LLaRA Model Card

2

5 commits

1 linked in READMEs

updated Jul 15, 2024

See the code

README


LLaRA Model Card

This model is released with paper LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

Xiang Li1, Cristina Mata1, Jongwoo Park1, Kumara Kahatapitiya1, Yoo Sung Jang1, Jinghuan Shang1, Kanchana Ranasinghe1, Ryan Burgert1, Mu Cai2, Yong Jae Lee2, and Michael S. Ryoo1

1Stony Brook University 2University of Wisconsin-Madison

Model details

Model type: LLaRA is an open-source visuomotor policy trained by fine-tuning LLaVA-7b-v1.5 on instruction-following data D-inBC and 4 auxiliary datasets, converted from VIMA-Data. For the conversion code, please refer to convert_vima.ipynb

Model date: llava-1.5-7b-llara-D-inBC-Aux-B-VIMA-80k was trained in June 2024.

Paper or resources for more information: https://github.com/LostXine/LLaRA

Where to send questions or comments about the model: https://github.com/LostXine/LLaRA/issues

Intended use

Primary intended uses: The primary use of LLaRA is research on large multimodal models for robotics.

Primary intended users: The primary intended users of the model are researchers and hobbyists in robotics, computer vision, natural language processing, machine learning, and artificial intelligence.

image-text-to-text
llara
llava
robotics
safetensors
text-generation
transformers

Contributors

variante

5 commits