ll-13/FIT-RS

Dataset

FIT-RS is a fine-grained remote sensing instruction tuning dataset, which contains 1,800,851 high-quality instruction samples covering various vision-language comprehension tasks. FIT-RS aims to enhance the fine-grained comprehension ability of Remote Sensing Large Multi-Modal Models (RSLMMs), specifically their ability to understand semantic relationships among objects in complex remote sensing scenes. The GitHub Repository is https://github.com/Luo-Z13/SkySenseGPT. Please refer to it for evalu

15

58 commits

2 linked in READMEs

updated Dec 11, 2024

See the code

README

FIT-RS is a fine-grained remote sensing instruction tuning dataset, which contains 1,800,851 high-quality instruction samples covering various vision-language comprehension tasks. FIT-RS aims to enhance the fine-grained comprehension ability of Remote Sensing Large Multi-Modal Models (RSLMMs), specifically their ability to understand semantic relationships among objects in complex remote sensing scenes. The GitHub Repository is https://github.com/Luo-Z13/SkySenseGPT. Please refer to it for evaluation and other details.

instruction-tuning
remote sensing
vision-language

ll-13/FIT-RS

Dataset

FIT-RS is a fine-grained remote sensing instruction tuning dataset, which contains 1,800,851 high-quality instruction samples covering various vision-language comprehension tasks. FIT-RS aims to enhance the fine-grained comprehension ability of Remote Sensing Large Multi-Modal Models (RSLMMs), specifically their ability to understand semantic relationships among objects in complex remote sensing scenes. The GitHub Repository is https://github.com/Luo-Z13/SkySenseGPT. Please refer to it for evalu

15

58 commits

2 linked in READMEs

updated Dec 11, 2024

See the code

README

FIT-RS is a fine-grained remote sensing instruction tuning dataset, which contains 1,800,851 high-quality instruction samples covering various vision-language comprehension tasks. FIT-RS aims to enhance the fine-grained comprehension ability of Remote Sensing Large Multi-Modal Models (RSLMMs), specifically their ability to understand semantic relationships among objects in complex remote sensing scenes. The GitHub Repository is https://github.com/Luo-Z13/SkySenseGPT. Please refer to it for evaluation and other details.

instruction-tuning
remote sensing
vision-language