[MobiSys 2026] A large-scale, multimodal dataset and benchmark for Human Action Recognition, Understanding and Reasoning
Python
305
25 commits
updated Oct 5, 2026
One recording, seven time-aligned streams: RGB · IR · Thermal · Depth · mmWave Radar · Skeleton · IMU
📥 Get the data · 🛰️ Explore it in your browser · 🏆 CUHK-X Challenge @ UbiComp/ISWC 2026 · Grand Finals Oct 12
CUHK-X is a comprehensive multimodal dataset containing 64,267 samples across seven modalities designed for human activity recognition, understanding, and reasoning. It addresses critical gaps in existing HAR datasets by providing synchronized multimodal sensor data with detailed annotations for complex reasoning tasks.
| Path | What it is |
|---|---|
SM/ | Small-model HAR baselines for RGB, IMU, mmWave radar and skeleton (PyTorch training and evaluation pipelines). |
LM/ | CUHK-X-VLM: large-model benchmark code for HAU and HARn tasks (action selection, captioning, emotion analysis, sequential reordering, next-action prediction) with QwenVL, InternVL, Video-LLaVA and Video-Chat. |
docs/ | Project website and the in-browser Observatory with curated depth / IR / thermal clips. |
scripts/ | Asset build scripts for the Observatory and the README banner. |
Dataset files are distributed separately under a Data Use Agreement. Request access on the dataset portal.
CITATION.cff.The CUHK-X Multimodal Human Activity Challenge is an official competition of UbiComp/ISWC 2026 and the first RGB-free international HAR competition. It runs on Kaggle in two parallel tracks built on CUHK-X, each with an independent USD $10K prize pool (USD $20K total). The baselines in this repository are the reference implementations for both tracks.
| Track | Task | Modalities | Constraints | Kaggle | Baselines here |
|---|---|---|---|---|---|
| Small Model | 40-class cross-subject HAR | Depth · IMU · mmWave radar · skeleton · IR · thermal | CNN / RNN / Transformer only · ≤ 100 MB · no large pretrained backbones, closed-source APIs or LLMs | Small Model Track | SM/ |
| Large Model | VQA for action understanding (HAU) and reasoning (HARn) | Depth · thermal · IR · skeleton · IMU · mmWave radar | No parameter limit · LVLMs, closed-source APIs and LLM pseudo-labeling allowed | Large Model Track | LM/ |
Both tracks train on users 1–9 and 16–24 and are scored cross-subject on Kaggle test users 10–11 and 25–26; finalists are additionally evaluated on organizer-held private data. RGB is excluded from the challenge.
Timeline
| Stage | Date (UTC) | Status |
|---|---|---|
| Kaggle competition open | Jun 20 – Sep 15, 2026 | ✅ Closed |
| Submission package due (Top 15 per track) | Sep 18, 2026 | ✅ Closed |
| Verification & private-data evaluation | Sep 19 – 30, 2026 | Scheduled period ended |
| Technical report deadline · scheduled Top 6 announcement | Oct 1, 2026 | Scheduled date passed; see organizer updates |
| Grand Finals @ UbiComp/ISWC 2026, Shanghai | Oct 12, 2026 | 📅 Upcoming |
Grand Finals. Yangtze River Hall, 5F, Shanghai International Convention Center. Small Model Track 09:45–12:30 and Large Model Track 14:30–17:20 (Shanghai time); each team gives a 5-minute talk followed by 5 minutes of Q&A, with remote participation via Zoom. Finalist scores combine the Kaggle private leaderboard (20%), organizer private-data evaluation (30%), reproducibility (10%), technical report (20%), presentation (10%) and model efficiency (10%). Finalist solutions are open-sourced under Apache 2.0 within 30 days of the finals.
Challenge data (as published on the challenge site; RGB excluded, test labels withheld). Small Model Track: Hugging Face · Google Drive · Baidu Netdisk. Large Model Track: Hugging Face · Google Drive · Baidu Netdisk.
Full rules, prizes and certificate tiers: UbiComp/ISWC 2026 competition page · challenge website · questions to cuhkx.competition@gmail.com.
The dataset is organized into two main components:
Objective: Traditional action classification across modalities
Objective: Comprehend actions through perceptual and contextual integration
Sub-tasks:
Objective: Infer intentions and causal relationships in action sequences
| Modality | Accuracy | F1 score | Precision | Recall |
|---|---|---|---|---|
| RGB | 90.89% | 91.28% | 92.24% | 91.02% |
| Depth | 90.46% | 90.93% | 91.76% | 90.75% |
| IR | 90.22% | 90.46% | 91.53% | 89.94% |
| Thermal | 92.57% | 93.36% | 93.54% | 93.50% |
| Radar | 46.63% | 44.53% | 48.29% | 46.63% |
| IMU | 45.52% | 38.32% | 40.84% | 38.00% |
| Skeleton | 79.08% | 84.17% | 91.46% | 79.08% |
| Model | Captioning(BLEU-1) | Emotion Analysis(Accuracy) | Sequential Reordering(Accuracy) |
|---|---|---|---|
| QwenVL-7B | 18.04% | 55.03% | 60.00% |
| VLLaVA-7B | 12.86% | 73.34% | 5.29% |
| InternVL-8B | 0.72% | 31.35% | 74.03% |
CUHK-X aims to advance research in:
If you use CUHK-X in your research, please cite our paper:
@inproceedings{10.1145/3745756.3809209,
author = {Jiang, Siyang and Yuan, Mu and Ji, Xiang and Yang, Bufang and Liu, Zeyu and Xu, Lilin and Li, Yang and He, Yuting and Dong, Liran and Lu, Wenrui and Yan, Zhenyu and Jiang, Xiaofan and Gao, Wei and Chen, Hongkai and Xing, Guoliang},
title = {A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning},
booktitle = {Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services},
series = {MobiSys '26},
year = {2026},
pages = {352--370},
numpages = {19},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
location = {University of Cambridge, Cambridge, United Kingdom},
isbn = {9798400720277},
doi = {10.1145/3745756.3809209},
url = {https://dl.acm.org/doi/epdf/10.1145/3745756.3809209},
keywords = {human action understanding, large language models, datasets}
}
A machine-readable version lives in CITATION.cff; GitHub's "Cite this repository" button uses it.
For dataset access, questions, or collaborations:
Code is released under the MIT License. The dataset is available for non-commercial research under a Data Use Agreement (DUA) and is not redistributable. Our derived annotations/splits are released under CC BY 4.0.
Note: This dataset is designed for research and educational purposes. Please ensure compliance with your institution's ethics guidelines when using human activity data.
We obtained approval from an Institutional Review Board (IRB) to conduct this study and collect data from human subjects.
[MobiSys 2026] A large-scale, multimodal dataset and benchmark for Human Action Recognition, Understanding and Reasoning
Python
305
25 commits
updated Oct 5, 2026
One recording, seven time-aligned streams: RGB · IR · Thermal · Depth · mmWave Radar · Skeleton · IMU
📥 Get the data · 🛰️ Explore it in your browser · 🏆 CUHK-X Challenge @ UbiComp/ISWC 2026 · Grand Finals Oct 12
CUHK-X is a comprehensive multimodal dataset containing 64,267 samples across seven modalities designed for human activity recognition, understanding, and reasoning. It addresses critical gaps in existing HAR datasets by providing synchronized multimodal sensor data with detailed annotations for complex reasoning tasks.
| Path | What it is |
|---|---|
SM/ | Small-model HAR baselines for RGB, IMU, mmWave radar and skeleton (PyTorch training and evaluation pipelines). |
LM/ | CUHK-X-VLM: large-model benchmark code for HAU and HARn tasks (action selection, captioning, emotion analysis, sequential reordering, next-action prediction) with QwenVL, InternVL, Video-LLaVA and Video-Chat. |
docs/ | Project website and the in-browser Observatory with curated depth / IR / thermal clips. |
scripts/ | Asset build scripts for the Observatory and the README banner. |
Dataset files are distributed separately under a Data Use Agreement. Request access on the dataset portal.
CITATION.cff.The CUHK-X Multimodal Human Activity Challenge is an official competition of UbiComp/ISWC 2026 and the first RGB-free international HAR competition. It runs on Kaggle in two parallel tracks built on CUHK-X, each with an independent USD $10K prize pool (USD $20K total). The baselines in this repository are the reference implementations for both tracks.
| Track | Task | Modalities | Constraints | Kaggle | Baselines here |
|---|---|---|---|---|---|
| Small Model | 40-class cross-subject HAR | Depth · IMU · mmWave radar · skeleton · IR · thermal | CNN / RNN / Transformer only · ≤ 100 MB · no large pretrained backbones, closed-source APIs or LLMs | Small Model Track | SM/ |
| Large Model | VQA for action understanding (HAU) and reasoning (HARn) | Depth · thermal · IR · skeleton · IMU · mmWave radar | No parameter limit · LVLMs, closed-source APIs and LLM pseudo-labeling allowed | Large Model Track | LM/ |
Both tracks train on users 1–9 and 16–24 and are scored cross-subject on Kaggle test users 10–11 and 25–26; finalists are additionally evaluated on organizer-held private data. RGB is excluded from the challenge.
Timeline
| Stage | Date (UTC) | Status |
|---|---|---|
| Kaggle competition open | Jun 20 – Sep 15, 2026 | ✅ Closed |
| Submission package due (Top 15 per track) | Sep 18, 2026 | ✅ Closed |
| Verification & private-data evaluation | Sep 19 – 30, 2026 | Scheduled period ended |
| Technical report deadline · scheduled Top 6 announcement | Oct 1, 2026 | Scheduled date passed; see organizer updates |
| Grand Finals @ UbiComp/ISWC 2026, Shanghai | Oct 12, 2026 | 📅 Upcoming |
Grand Finals. Yangtze River Hall, 5F, Shanghai International Convention Center. Small Model Track 09:45–12:30 and Large Model Track 14:30–17:20 (Shanghai time); each team gives a 5-minute talk followed by 5 minutes of Q&A, with remote participation via Zoom. Finalist scores combine the Kaggle private leaderboard (20%), organizer private-data evaluation (30%), reproducibility (10%), technical report (20%), presentation (10%) and model efficiency (10%). Finalist solutions are open-sourced under Apache 2.0 within 30 days of the finals.
Challenge data (as published on the challenge site; RGB excluded, test labels withheld). Small Model Track: Hugging Face · Google Drive · Baidu Netdisk. Large Model Track: Hugging Face · Google Drive · Baidu Netdisk.
Full rules, prizes and certificate tiers: UbiComp/ISWC 2026 competition page · challenge website · questions to cuhkx.competition@gmail.com.
The dataset is organized into two main components:
Objective: Traditional action classification across modalities
Objective: Comprehend actions through perceptual and contextual integration
Sub-tasks:
Objective: Infer intentions and causal relationships in action sequences
| Modality | Accuracy | F1 score | Precision | Recall |
|---|---|---|---|---|
| RGB | 90.89% | 91.28% | 92.24% | 91.02% |
| Depth | 90.46% | 90.93% | 91.76% | 90.75% |
| IR | 90.22% | 90.46% | 91.53% | 89.94% |
| Thermal | 92.57% | 93.36% | 93.54% | 93.50% |
| Radar | 46.63% | 44.53% | 48.29% | 46.63% |
| IMU | 45.52% | 38.32% | 40.84% | 38.00% |
| Skeleton | 79.08% | 84.17% | 91.46% | 79.08% |
| Model | Captioning(BLEU-1) | Emotion Analysis(Accuracy) | Sequential Reordering(Accuracy) |
|---|---|---|---|
| QwenVL-7B | 18.04% | 55.03% | 60.00% |
| VLLaVA-7B | 12.86% | 73.34% | 5.29% |
| InternVL-8B | 0.72% | 31.35% | 74.03% |
CUHK-X aims to advance research in:
If you use CUHK-X in your research, please cite our paper:
@inproceedings{10.1145/3745756.3809209,
author = {Jiang, Siyang and Yuan, Mu and Ji, Xiang and Yang, Bufang and Liu, Zeyu and Xu, Lilin and Li, Yang and He, Yuting and Dong, Liran and Lu, Wenrui and Yan, Zhenyu and Jiang, Xiaofan and Gao, Wei and Chen, Hongkai and Xing, Guoliang},
title = {A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning},
booktitle = {Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services},
series = {MobiSys '26},
year = {2026},
pages = {352--370},
numpages = {19},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
location = {University of Cambridge, Cambridge, United Kingdom},
isbn = {9798400720277},
doi = {10.1145/3745756.3809209},
url = {https://dl.acm.org/doi/epdf/10.1145/3745756.3809209},
keywords = {human action understanding, large language models, datasets}
}
A machine-readable version lives in CITATION.cff; GitHub's "Cite this repository" button uses it.
For dataset access, questions, or collaborations:
Code is released under the MIT License. The dataset is available for non-commercial research under a Data Use Agreement (DUA) and is not redistributable. Our derived annotations/splits are released under CC BY 4.0.
Note: This dataset is designed for research and educational purposes. Please ensure compliance with your institution's ethics guidelines when using human activity data.
We obtained approval from an Institutional Review Board (IRB) to conduct this study and collect data from human subjects.