9 repos
Vision-language models and frameworks designed to understand and reason about spatial relationships, geometry, and 3D scene understanding. This cluster covers VLM architectures optimized for spatial intelligence tasks, along with curated resources and benchmarks for evaluating spatial reasoning capabilities in multimodal AI systems. The core focus is enabling language models to process visual information with geometric and spatial awareness.