An AI dataset index covering major research areas such as NLP, computer vision, multimodal learning, speech, audio, music understanding, time series, graph learning, recommender systems, retrieval, LLMs, agents, computer use, mobile UI, robotics, embodied AI, autonomous driving, remote sensing, scientific AI, medical AI, and related domains.
Unknown.| Category | Focus | English | 中文 | Count |
|---|---|---|---|---|
| NLP | Text datasets for NER, QA, summarization, reasoning, classification, multilingual transfer, and language understanding. | Open | 中文 | 29 |
| CV | Image datasets for classification, detection, segmentation, grounding, scene understanding, RGB-D perception, and fine-grained recognition. | Open | 中文 | 43 |
| Video-3D | Video, egocentric, point-cloud, 3D scene, and shape datasets. | Open | 中文 | 21 |
| Autonomous-Driving | Driving perception, 3D detection, tracking, forecasting, mapping, and cooperative sensing datasets. | Open | 中文 | 24 |
| Remote-Sensing | Satellite, aerial, overhead, multispectral, SAR, and geospatial mapping datasets. | Open | 中文 | 23 |
| Multimodal | Vision-language, VQA, chart and OCR-grounded reasoning, image-text alignment, and multimodal instruction datasets. | Open | 中文 | 27 |
| Speech-Audio | ASR, speech translation, speaker or language ID, speech emotion, and speech generation datasets. | Open | 中文 | 25 |
| Audio-Understanding | Sound events, audio captioning, audio-language learning, and non-music audio foundation-model evaluation datasets. | Open | 中文 | 17 |
| Music-Audio | Music MIR, transcription, source separation, tagging, singing voice, and symbolic-audio learning datasets. | Open | 中文 | 15 |
| Time-Series | Forecasting, classification, anomaly detection, clinical time series, and spatiotemporal sequence datasets. | Open | 中文 | 24 |
| Document-AI | OCR, layout analysis, forms, receipts, tables, charts, and document QA datasets. | Open | 中文 | 21 |
| Code | Code generation, repair, execution, repository understanding, and software engineering agent datasets. | Open | 中文 | 21 |
| Search-Retrieval | Embedding, retrieval, reranking, multilingual IR, RAG grounding, and multimodal or audio-text retrieval datasets. | Open | 中文 | 23 |
| Graph-Learning | Node, link, graph-level, molecular, knowledge, heterogeneous, and temporal graph datasets. | Open | 中文 | 22 |
| Recommender-Systems | Collaborative filtering, ranking, CTR, news, bandit, and industrial recommendation datasets. | Open | 中文 | 22 |
| LLM | Pretraining corpora, instruction tuning, synthetic supervision, preference data, and alignment datasets. | Open | 中文 | 26 |
| LLM-Evals | Instruction following, reasoning, truthfulness, chat, and long-context benchmark datasets for large language models. | Open | 中文 | 28 |
| Agent | Tool use, software engineering, long-context memory, and general interactive agent datasets. | Open | 中文 | 16 |
| Computer-Use | Browser, desktop, and cross-platform GUI-grounded datasets and benchmarks for computer-use agents. | Open | 中文 | 18 |
| Mobile-UI | Mobile UI understanding, screen grounding, and Android agent datasets and benchmarks. | Open | 中文 | 18 |
| Robotics-RL | Offline RL, robot-learning benchmarks, simulation environments, and control-oriented policy suites. | Open | 中文 | 14 |
| Robot-Manipulation | Real-world robot manipulation datasets, teleoperation corpora, and visuomotor trajectory collections. | Open | 中文 | 14 |
| Embodied-AI | Embodied navigation, instruction following, egocentric perception, and language-conditioned behavior datasets. | Open | 中文 | 21 |
| Scientific-AI | Molecule, protein, material, reaction, and scientific literature datasets. | Open | 中文 | 23 |
| Medical-AI | Clinical, biomedical NLP, radiology, pathology, medical imaging, and healthcare datasets. | Open | 中文 | 26 |
| Finance-Legal | Finance, regulation, legal reasoning, extraction, filing analysis, contract understanding, and compliance datasets. | Open | 中文 | 20 |
| Benchmarks | Cross-domain evaluation suites for NLP, CV, multimodal systems, retrieval, code, agents, and general platforms. | Open | 中文 | 25 |
| Data-Portals | Dataset registries, search portals, and discovery hubs for AI data and benchmarks. | Open | 中文 | 10 |
| Safety-Evals | Safety eval suites, risk taxonomies, red-teaming collections, and agent safety resources. | Open | 中文 | 20 |
35 commits
An AI dataset index covering major research areas such as NLP, computer vision, multimodal learning, speech, audio, music understanding, time series, graph learning, recommender systems, retrieval, LLMs, agents, computer use, mobile UI, robotics, embodied AI, autonomous driving, remote sensing, scientific AI, medical AI, and related domains.
Unknown.| Category | Focus | English | 中文 | Count |
|---|---|---|---|---|
| NLP | Text datasets for NER, QA, summarization, reasoning, classification, multilingual transfer, and language understanding. | Open | 中文 | 29 |
| CV | Image datasets for classification, detection, segmentation, grounding, scene understanding, RGB-D perception, and fine-grained recognition. | Open | 中文 | 43 |
| Video-3D | Video, egocentric, point-cloud, 3D scene, and shape datasets. | Open | 中文 | 21 |
| Autonomous-Driving | Driving perception, 3D detection, tracking, forecasting, mapping, and cooperative sensing datasets. | Open | 中文 | 24 |
| Remote-Sensing | Satellite, aerial, overhead, multispectral, SAR, and geospatial mapping datasets. | Open | 中文 | 23 |
| Multimodal | Vision-language, VQA, chart and OCR-grounded reasoning, image-text alignment, and multimodal instruction datasets. | Open | 中文 | 27 |
| Speech-Audio | ASR, speech translation, speaker or language ID, speech emotion, and speech generation datasets. | Open | 中文 | 25 |
| Audio-Understanding | Sound events, audio captioning, audio-language learning, and non-music audio foundation-model evaluation datasets. | Open | 中文 | 17 |
| Music-Audio | Music MIR, transcription, source separation, tagging, singing voice, and symbolic-audio learning datasets. | Open | 中文 | 15 |
| Time-Series | Forecasting, classification, anomaly detection, clinical time series, and spatiotemporal sequence datasets. | Open | 中文 | 24 |
| Document-AI | OCR, layout analysis, forms, receipts, tables, charts, and document QA datasets. | Open | 中文 | 21 |
| Code | Code generation, repair, execution, repository understanding, and software engineering agent datasets. | Open | 中文 | 21 |
| Search-Retrieval | Embedding, retrieval, reranking, multilingual IR, RAG grounding, and multimodal or audio-text retrieval datasets. | Open | 中文 | 23 |
| Graph-Learning | Node, link, graph-level, molecular, knowledge, heterogeneous, and temporal graph datasets. | Open | 中文 | 22 |
| Recommender-Systems | Collaborative filtering, ranking, CTR, news, bandit, and industrial recommendation datasets. | Open | 中文 | 22 |
| LLM | Pretraining corpora, instruction tuning, synthetic supervision, preference data, and alignment datasets. | Open | 中文 | 26 |
| LLM-Evals | Instruction following, reasoning, truthfulness, chat, and long-context benchmark datasets for large language models. | Open | 中文 | 28 |
| Agent | Tool use, software engineering, long-context memory, and general interactive agent datasets. | Open | 中文 | 16 |
| Computer-Use | Browser, desktop, and cross-platform GUI-grounded datasets and benchmarks for computer-use agents. | Open | 中文 | 18 |
| Mobile-UI | Mobile UI understanding, screen grounding, and Android agent datasets and benchmarks. | Open | 中文 | 18 |
| Robotics-RL | Offline RL, robot-learning benchmarks, simulation environments, and control-oriented policy suites. | Open | 中文 | 14 |
| Robot-Manipulation | Real-world robot manipulation datasets, teleoperation corpora, and visuomotor trajectory collections. | Open | 中文 | 14 |
| Embodied-AI | Embodied navigation, instruction following, egocentric perception, and language-conditioned behavior datasets. | Open | 中文 | 21 |
| Scientific-AI | Molecule, protein, material, reaction, and scientific literature datasets. | Open | 中文 | 23 |
| Medical-AI | Clinical, biomedical NLP, radiology, pathology, medical imaging, and healthcare datasets. | Open | 中文 | 26 |
| Finance-Legal | Finance, regulation, legal reasoning, extraction, filing analysis, contract understanding, and compliance datasets. | Open | 中文 | 20 |
| Benchmarks | Cross-domain evaluation suites for NLP, CV, multimodal systems, retrieval, code, agents, and general platforms. | Open | 中文 | 25 |
| Data-Portals | Dataset registries, search portals, and discovery hubs for AI data and benchmarks. | Open | 中文 | 10 |
| Safety-Evals | Safety eval suites, risk taxonomies, red-teaming collections, and agent safety resources. | Open | 中文 | 20 |
35 commits