13 repos
Resources for building, organizing, and curating large-scale datasets and knowledge graphs, with emphasis on structured data collection, semantic organization, and dataset documentation. The cluster spans tools for dataset assembly (like wiki25 and UEGPT-Datasets for large-scale data harvesting), knowledge representation systems (Hive), and data quality/linking infrastructure (LIDQ, ALDEN). Projects here address the practical challenge of preparing high-quality, well-documented datasets for machine learning and semantic applications.