Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resources
2,263
478 commits
updated Apr 16, 2026
Here is a curated list of papers about 3D-Related Tasks empowered by Large Language Models (LLMs). It contains various tasks including 3D understanding, reasoning, generation, and embodied agents. Also, we include other Foundation Models (CLIP, SAM) for the whole picture of this area.
This is an active repository, you can watch for following the latest advances. If you find it useful, please kindly star ⭐ this repo and cite the paper.
| Date | Keywords | Institute (first) | Paper | Publication | Others |
|---|---|---|---|---|---|
| 2025-11-07 | Omni-View | PKU | Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images | ICLR 2026 | github |
| 2025-08-16 | UniUGG | FDU | UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding | ICLR 2026 | github |
| Date | keywords | Institute (first) | Paper | Publication | Others |
|---|---|---|---|---|---|
| 2025-12-15 | RoboTracer | BUAA | RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics | Arxiv | project |
| 2025-06-11 | SceneCOT | BIGAI | SceneCOT: Eliciting Chain-of-Thought Reasoning in 3D Scenes | Arxiv | project |
| 2025-06-04 | RoboRefer | BUAA | RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics | Arxiv | project |
| 2024-09-08 | MSR3D | BIGAI | Multi-modal Situated Reasoning in 3D Scenes | NeurIPS '24 | project |
| 2023-5-20 | 3D-CLR | UCLA | 3D Concept Learning and Reasoning from Multi-View Images | CVPR '23 | github |
| - | Transcribe3D | TTI, Chicago | Transcribe3D: Grounding LLMs Using Transcribed Information for 3D Referential Reasoning with Self-Corrected Finetuning | CoRL '23 | github |
Your contributions are always welcome!
I will keep some pull requests open if I'm not sure if they are awesome for 3D LLMs, you could vote for them by adding 👍 to them.
If you have any questions about this opinionated list, please get in touch at xianzheng@robots.ox.ac.uk or Wechat ID: mxz1997112.
If you find this repository useful, please consider citing this paper:
@misc{ma2024llmsstep3dworld,
title={When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models},
author={Xianzheng Ma and Yash Bhalgat and Brandon Smart and Shuai Chen and Xinghui Li and Jian Ding and Jindong Gu and Dave Zhenyu Chen and Songyou Peng and Jia-Wang Bian and Philip H Torr and Marc Pollefeys and Matthias Nießner and Ian D Reid and Angel X. Chang and Iro Laina and Victor Adrian Prisacariu},
year={2024},
journal={arXiv preprint arXiv:2405.10255},
}
This repo is inspired by Awesome-LLM
(top 30 of 98)
Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resources
2,263
478 commits
updated Apr 16, 2026
Here is a curated list of papers about 3D-Related Tasks empowered by Large Language Models (LLMs). It contains various tasks including 3D understanding, reasoning, generation, and embodied agents. Also, we include other Foundation Models (CLIP, SAM) for the whole picture of this area.
This is an active repository, you can watch for following the latest advances. If you find it useful, please kindly star ⭐ this repo and cite the paper.
| Date | Keywords | Institute (first) | Paper | Publication | Others |
|---|---|---|---|---|---|
| 2025-11-07 | Omni-View | PKU | Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images | ICLR 2026 | github |
| 2025-08-16 | UniUGG | FDU | UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding | ICLR 2026 | github |
| Date | keywords | Institute (first) | Paper | Publication | Others |
|---|---|---|---|---|---|
| 2025-12-15 | RoboTracer | BUAA | RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics | Arxiv | project |
| 2025-06-11 | SceneCOT | BIGAI | SceneCOT: Eliciting Chain-of-Thought Reasoning in 3D Scenes | Arxiv | project |
| 2025-06-04 | RoboRefer | BUAA | RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics | Arxiv | project |
| 2024-09-08 | MSR3D | BIGAI | Multi-modal Situated Reasoning in 3D Scenes | NeurIPS '24 | project |
| 2023-5-20 | 3D-CLR | UCLA | 3D Concept Learning and Reasoning from Multi-View Images | CVPR '23 | github |
| - | Transcribe3D | TTI, Chicago | Transcribe3D: Grounding LLMs Using Transcribed Information for 3D Referential Reasoning with Self-Corrected Finetuning | CoRL '23 | github |
Your contributions are always welcome!
I will keep some pull requests open if I'm not sure if they are awesome for 3D LLMs, you could vote for them by adding 👍 to them.
If you have any questions about this opinionated list, please get in touch at xianzheng@robots.ox.ac.uk or Wechat ID: mxz1997112.
If you find this repository useful, please consider citing this paper:
@misc{ma2024llmsstep3dworld,
title={When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models},
author={Xianzheng Ma and Yash Bhalgat and Brandon Smart and Shuai Chen and Xinghui Li and Jian Ding and Jindong Gu and Dave Zhenyu Chen and Songyou Peng and Jia-Wang Bian and Philip H Torr and Marc Pollefeys and Matthias Nießner and Ian D Reid and Angel X. Chang and Iro Laina and Victor Adrian Prisacariu},
year={2024},
journal={arXiv preprint arXiv:2405.10255},
}
This repo is inspired by Awesome-LLM
(top 30 of 98)