DAMO-NLP-SG/VideoRefer-7B

Model

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

5

5 commits

2 linked in READMEs

updated Dec 31, 2024

See the code

README

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

If you like our project, please give us a star ⭐ on Github for the latest update.

🌏 Model Zoo

πŸ“‘ Citation

If you find VideoRefer Suite useful for your research and applications, please cite using this BibTeX:

@article{yuan2024videorefersuite,
  title = {VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM},
  author = {Yuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng, Boqiang Zhang, Long Li, Xin Li, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing},
  journal={arXiv},
  year={2024},
  url = {}
}
endpoints_compatible
large video-language model
multimodal large language model
safetensors
text-generation
transformers
videorefer_qwen2
visual-question-answering

DAMO-NLP-SG/VideoRefer-7B

Model

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

5

5 commits

2 linked in READMEs

updated Dec 31, 2024

See the code

README

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

If you like our project, please give us a star ⭐ on Github for the latest update.

🌏 Model Zoo

πŸ“‘ Citation

If you find VideoRefer Suite useful for your research and applications, please cite using this BibTeX:

@article{yuan2024videorefersuite,
  title = {VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM},
  author = {Yuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng, Boqiang Zhang, Long Li, Xin Li, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing},
  journal={arXiv},
  year={2024},
  url = {}
}
endpoints_compatible
large video-language model
multimodal large language model
safetensors
text-generation
transformers
videorefer_qwen2
visual-question-answering