63 repos across 3 sub-areas
Physical AI systems that combine visual perception and language understanding to enable robots to perform complex tasks and manipulation. The cluster encompasses vision-language-action (VLA) models, robotic learning frameworks, and datasets for training embodied AI agents to understand and execute human instructions in physical environments. Central repositories include large language models adapted for robotics (GigaBrain), reinforcement learning frameworks for embodied control (RLDX), and task-specific fine-tuned variants for simulation environments and real robot systems.
Vision Language Action Models for Robotics
30 repos
Vision Language Action (VLA) models that combine visual perception, language understanding, and robotic control into unified transformer-based architectures. These repositories implement and extend VLA frameworks for robot learning, instruction following, and multi-task manipulation, often leveraging pre-trained transformers and safetensors for efficient model storage and deployment. The cluster includes research implementations, benchmark environments, and real-world robot applications ranging from single-arm manipulation to multi-robot coordination.
Robotic Vision-Language Action Models
23 repos
Vision-language-action (VLA) models and learning frameworks that enable robots to understand and execute physical tasks from visual observations and language instructions. The cluster covers foundation models for robotic manipulation and navigation, reinforcement learning approaches for skill acquisition, and simulation environments for training embodied AI agents. Key themes include behavior cloning from demonstrations, world model learning, and real-to-sim transfer for continuous control in kitchen and household manipulation tasks.
Vision-Language-Action Models for Robotics
10 repos
Vision-language-action (VLA) models that enable robots to understand and act on natural language instructions by integrating visual perception with motor control. This cluster focuses on model architectures and implementations for embodied AI, including optimized variants like SmolVLA and BitVLA designed for efficient deployment on edge devices and robotics platforms like Jetson. Repositories here span model checkpoints, inference frameworks, and robot task implementations that bridge language understanding with robotic manipulation and control.