7 repos
Qi-Zhangyang/GPT4Scene
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
539
18 commits
InternRobotics/G2VLM
[CVPR 2026] G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and…
354
16 commits
VITA-Group/VLM-3R
[CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
447
26 commits
AIGeeksGroup/3D-R1
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
415
7 commits
Qi-Zhangyang/GPT4Scene-and-VLN-R1
LiyaoTang/g2vlm
No description
0
4 commits
ZCMax/LLaVA-3D
[ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D…
388
12 commits