12 repos
Tools, frameworks, and utilities for managing, processing, and organizing large-scale datasets and data pipelines. The cluster includes systems for dataset curation, distributed data handling, and infrastructure for machine learning workflows. Central repositories like LIDQ, ALDEN, and the Hive-related projects suggest a focus on practical data engineering challenges, while the diverse set of contributions indicates work across multiple data management paradigms and use cases.