lyirs/AIDataset

An AI dataset index covering major research areas

31

35 commits

updated Apr 25, 2026

See the code

README

AIDataset

An AI dataset index covering major research areas such as NLP, computer vision, multimodal learning, speech, audio, music understanding, time series, graph learning, recommender systems, retrieval, LLMs, agents, computer use, mobile UI, robotics, embodied AI, autonomous driving, remote sensing, scientific AI, medical AI, and related domains.

中文说明

Snapshot

  • 29 topical directories
  • 636 primary entries
  • Links checked on 2026-04-26

Inclusion Rules

  • Public datasets, benchmark suites, or discovery portals with clear value for training, evaluation, or research navigation.
  • Priority is given to datasets that remain common in top-conference papers, tutorials, and baseline comparison tables.
  • Official sites, official GitHub repositories, Hugging Face dataset cards, or trusted public portals.
  • Licenses are written as published; if the source is unclear, the table says Unknown.
  • Some benchmark and safety resources are registries or taxonomies rather than raw downloadable datasets. Those are explicitly labeled in their section notes.
  • This repository is an index, not a mirror. No dataset files are redistributed here.

Category Index

CategoryFocusEnglish中文Count
NLPText datasets for NER, QA, summarization, reasoning, classification, multilingual transfer, and language understanding.Open中文29
CVImage datasets for classification, detection, segmentation, grounding, scene understanding, RGB-D perception, and fine-grained recognition.Open中文43
Video-3DVideo, egocentric, point-cloud, 3D scene, and shape datasets.Open中文21
Autonomous-DrivingDriving perception, 3D detection, tracking, forecasting, mapping, and cooperative sensing datasets.Open中文24
Remote-SensingSatellite, aerial, overhead, multispectral, SAR, and geospatial mapping datasets.Open中文23
MultimodalVision-language, VQA, chart and OCR-grounded reasoning, image-text alignment, and multimodal instruction datasets.Open中文27
Speech-AudioASR, speech translation, speaker or language ID, speech emotion, and speech generation datasets.Open中文25
Audio-UnderstandingSound events, audio captioning, audio-language learning, and non-music audio foundation-model evaluation datasets.Open中文17
Music-AudioMusic MIR, transcription, source separation, tagging, singing voice, and symbolic-audio learning datasets.Open中文15
Time-SeriesForecasting, classification, anomaly detection, clinical time series, and spatiotemporal sequence datasets.Open中文24
Document-AIOCR, layout analysis, forms, receipts, tables, charts, and document QA datasets.Open中文21
CodeCode generation, repair, execution, repository understanding, and software engineering agent datasets.Open中文21
Search-RetrievalEmbedding, retrieval, reranking, multilingual IR, RAG grounding, and multimodal or audio-text retrieval datasets.Open中文23
Graph-LearningNode, link, graph-level, molecular, knowledge, heterogeneous, and temporal graph datasets.Open中文22
Recommender-SystemsCollaborative filtering, ranking, CTR, news, bandit, and industrial recommendation datasets.Open中文22
LLMPretraining corpora, instruction tuning, synthetic supervision, preference data, and alignment datasets.Open中文26
LLM-EvalsInstruction following, reasoning, truthfulness, chat, and long-context benchmark datasets for large language models.Open中文28
AgentTool use, software engineering, long-context memory, and general interactive agent datasets.Open中文16
Computer-UseBrowser, desktop, and cross-platform GUI-grounded datasets and benchmarks for computer-use agents.Open中文18
Mobile-UIMobile UI understanding, screen grounding, and Android agent datasets and benchmarks.Open中文18
Robotics-RLOffline RL, robot-learning benchmarks, simulation environments, and control-oriented policy suites.Open中文14
Robot-ManipulationReal-world robot manipulation datasets, teleoperation corpora, and visuomotor trajectory collections.Open中文14
Embodied-AIEmbodied navigation, instruction following, egocentric perception, and language-conditioned behavior datasets.Open中文21
Scientific-AIMolecule, protein, material, reaction, and scientific literature datasets.Open中文23
Medical-AIClinical, biomedical NLP, radiology, pathology, medical imaging, and healthcare datasets.Open中文26
Finance-LegalFinance, regulation, legal reasoning, extraction, filing analysis, contract understanding, and compliance datasets.Open中文20
BenchmarksCross-domain evaluation suites for NLP, CV, multimodal systems, retrieval, code, agents, and general platforms.Open中文25
Data-PortalsDataset registries, search portals, and discovery hubs for AI data and benchmarks.Open中文10
Safety-EvalsSafety eval suites, risk taxonomies, red-teaming collections, and agent safety resources.Open中文20

Source Families

  • GitHub: official organization repositories, task repositories, benchmark suites, and maintainers.
  • Public portals: Hugging Face, OpenML, OpenDataLab, UCI, PhysioNet, OpenSLR, NIST, and research portals.
  • Official websites: challenge homepages, dataset landing pages, project sites, and leaderboard pages.

Contributors

lyirs

35 commits

lyirs/AIDataset

An AI dataset index covering major research areas

31

35 commits

updated Apr 25, 2026

See the code

README

AIDataset

An AI dataset index covering major research areas such as NLP, computer vision, multimodal learning, speech, audio, music understanding, time series, graph learning, recommender systems, retrieval, LLMs, agents, computer use, mobile UI, robotics, embodied AI, autonomous driving, remote sensing, scientific AI, medical AI, and related domains.

中文说明

Snapshot

  • 29 topical directories
  • 636 primary entries
  • Links checked on 2026-04-26

Inclusion Rules

  • Public datasets, benchmark suites, or discovery portals with clear value for training, evaluation, or research navigation.
  • Priority is given to datasets that remain common in top-conference papers, tutorials, and baseline comparison tables.
  • Official sites, official GitHub repositories, Hugging Face dataset cards, or trusted public portals.
  • Licenses are written as published; if the source is unclear, the table says Unknown.
  • Some benchmark and safety resources are registries or taxonomies rather than raw downloadable datasets. Those are explicitly labeled in their section notes.
  • This repository is an index, not a mirror. No dataset files are redistributed here.

Category Index

CategoryFocusEnglish中文Count
NLPText datasets for NER, QA, summarization, reasoning, classification, multilingual transfer, and language understanding.Open中文29
CVImage datasets for classification, detection, segmentation, grounding, scene understanding, RGB-D perception, and fine-grained recognition.Open中文43
Video-3DVideo, egocentric, point-cloud, 3D scene, and shape datasets.Open中文21
Autonomous-DrivingDriving perception, 3D detection, tracking, forecasting, mapping, and cooperative sensing datasets.Open中文24
Remote-SensingSatellite, aerial, overhead, multispectral, SAR, and geospatial mapping datasets.Open中文23
MultimodalVision-language, VQA, chart and OCR-grounded reasoning, image-text alignment, and multimodal instruction datasets.Open中文27
Speech-AudioASR, speech translation, speaker or language ID, speech emotion, and speech generation datasets.Open中文25
Audio-UnderstandingSound events, audio captioning, audio-language learning, and non-music audio foundation-model evaluation datasets.Open中文17
Music-AudioMusic MIR, transcription, source separation, tagging, singing voice, and symbolic-audio learning datasets.Open中文15
Time-SeriesForecasting, classification, anomaly detection, clinical time series, and spatiotemporal sequence datasets.Open中文24
Document-AIOCR, layout analysis, forms, receipts, tables, charts, and document QA datasets.Open中文21
CodeCode generation, repair, execution, repository understanding, and software engineering agent datasets.Open中文21
Search-RetrievalEmbedding, retrieval, reranking, multilingual IR, RAG grounding, and multimodal or audio-text retrieval datasets.Open中文23
Graph-LearningNode, link, graph-level, molecular, knowledge, heterogeneous, and temporal graph datasets.Open中文22
Recommender-SystemsCollaborative filtering, ranking, CTR, news, bandit, and industrial recommendation datasets.Open中文22
LLMPretraining corpora, instruction tuning, synthetic supervision, preference data, and alignment datasets.Open中文26
LLM-EvalsInstruction following, reasoning, truthfulness, chat, and long-context benchmark datasets for large language models.Open中文28
AgentTool use, software engineering, long-context memory, and general interactive agent datasets.Open中文16
Computer-UseBrowser, desktop, and cross-platform GUI-grounded datasets and benchmarks for computer-use agents.Open中文18
Mobile-UIMobile UI understanding, screen grounding, and Android agent datasets and benchmarks.Open中文18
Robotics-RLOffline RL, robot-learning benchmarks, simulation environments, and control-oriented policy suites.Open中文14
Robot-ManipulationReal-world robot manipulation datasets, teleoperation corpora, and visuomotor trajectory collections.Open中文14
Embodied-AIEmbodied navigation, instruction following, egocentric perception, and language-conditioned behavior datasets.Open中文21
Scientific-AIMolecule, protein, material, reaction, and scientific literature datasets.Open中文23
Medical-AIClinical, biomedical NLP, radiology, pathology, medical imaging, and healthcare datasets.Open中文26
Finance-LegalFinance, regulation, legal reasoning, extraction, filing analysis, contract understanding, and compliance datasets.Open中文20
BenchmarksCross-domain evaluation suites for NLP, CV, multimodal systems, retrieval, code, agents, and general platforms.Open中文25
Data-PortalsDataset registries, search portals, and discovery hubs for AI data and benchmarks.Open中文10
Safety-EvalsSafety eval suites, risk taxonomies, red-teaming collections, and agent safety resources.Open中文20

Source Families

  • GitHub: official organization repositories, task repositories, benchmark suites, and maintainers.
  • Public portals: Hugging Face, OpenML, OpenDataLab, UCI, PhysioNet, OpenSLR, NIST, and research portals.
  • Official websites: challenge homepages, dataset landing pages, project sites, and leaderboard pages.

Contributors

lyirs

35 commits