AntResearchNLP/ViLaSR

Model

18

stars

5

commits

1

repos using this model

1

linked in READMEs

Aug 11, 2025

updated

conversational
image-text-to-text
qwen2_5_vl
safetensors

README

This repository contains the ViLaSR-7B model as presented in Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.

Please refer to the code https://github.com/AntResearchNLP/ViLaSR.

@misc{wu2025reinforcingspatialreasoningvisionlanguage,
      title={Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing}, 
      author={Junfei Wu and Jian Guan and Kaituo Feng and Qiang Liu and Shu Wu and Liang Wang and Wei Wu and Tieniu Tan},
      year={2025},
      eprint={2506.09965},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2506.09965}, 
}

Contributors

Hyperwjf

5 commits

AntResearchNLP/ViLaSR

Model

18

stars

5

commits

1

repos using this model

1

linked in READMEs

Aug 11, 2025

updated

conversational
image-text-to-text
qwen2_5_vl
safetensors

README

This repository contains the ViLaSR-7B model as presented in Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.

Please refer to the code https://github.com/AntResearchNLP/ViLaSR.

@misc{wu2025reinforcingspatialreasoningvisionlanguage,
      title={Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing}, 
      author={Junfei Wu and Jian Guan and Kaituo Feng and Qiang Liu and Shu Wu and Liang Wang and Wei Wu and Tieniu Tan},
      year={2025},
      eprint={2506.09965},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2506.09965}, 
}

Contributors

Hyperwjf

5 commits