12 repos
Methods and models for processing, understanding, and generating 3D point cloud data and scene representations. This cluster includes approaches for 3D object detection, scene understanding, multi-modal learning that combines vision with language (as seen in chatbot-integrated models like MiniGPT-3D), and foundational representation learning on large-scale 3D datasets like Objaverse. Repositories here span model architectures for temporal and spatial 3D reasoning, often implemented in Python with deep learning frameworks.