TLDR: This is a collection of open-source video instruction tuning data used in
Cambrian-S's third training stage.
Cambrian-S-3M combines three video instruction datasets:
pip install -U "huggingface_hub[cli]==0.36.0"hf command should be available after installing huggingface_hubcd $PATH_TO_LOCAL_DATASET
# Download Cambrian-S-3M to local disk
hf download nyu-visionx/Cambrian-S-3M --repo-type dataset --local-dir Cambrian-S-3M
# Decompress all archives
cd Cambrian-S-3M
bash decompress.sh
See decompress.sh for details.
cd $PATH_TO_LOCAL_DATASET
# Download LLaVA-Video-178K
hf download lmms-lab/LLaVA-Video-178K --repo-type dataset --local-dir LLaVA-Video-178K
# Decompress all archives
cd LLaVA-Video-178K
find . -name "*.tar.gz" -exec tar -zxf {} \;
cd $PATH_TO_LOCAL_DATASET
# Download LLaVA-Hound (ShareGPTVideo)
hf download ShareGPTVideo/train_video_and_instruction --repo-type dataset --local-dir train_video_and_instruction
# Decompress all archives
cd train_video_and_instruction/train_300k
find . -name "*.tar.gz" -exec tar -zxf {} \;
Link the additional datasets into the Cambrian-S-3M directory:
cd $PATH_TO_LOCAL_DATASET/Cambrian-S-3M
# Link LLaVA-Video-178K files
ln -s $PATH_TO_LOCAL_DATASET/LLaVA-Video-178K/* ./
# Link LLaVA-Hound frames
mkdir -p shareVideoGPTV/frames
ln -s $PATH_TO_LOCAL_DATASET/train_video_and_instruction/train_300k ./shareVideoGPTV/frames/all_frames
Run the sanity check to ensure all data is properly set up:
cd $PATH_TO_LOCAL_DATASET/Cambrian-S-3M
python sanity_check.py
hf command: Install with pip install -U "huggingface_hub[cli]"$PATH_TO_LOCAL_DATASETIf you use this dataset, please cite:
@article{yang2025cambrians,
title={Cambrian-S: Towards Spatial Supersensing in Video},
author={Yang, Shusheng and Yang, Jihan and Huang, Pinzhi and Brown, Ellis and Yang, Zihao and Yu, Yue and Tong, Shengbang and Zheng, Zihan and Xu, Yifan and Wang, Muhan and Lu, Daohan and Fergus, Rob and LeCun, Yann and Fei-Fei, Li and Xie, Saining},
journal={arXiv preprint arXiv:2511.04670},
year={2025}
}
This dataset is released under the Apache 2.0 license.
500 commits
TLDR: This is a collection of open-source video instruction tuning data used in
Cambrian-S's third training stage.
Cambrian-S-3M combines three video instruction datasets:
pip install -U "huggingface_hub[cli]==0.36.0"hf command should be available after installing huggingface_hubcd $PATH_TO_LOCAL_DATASET
# Download Cambrian-S-3M to local disk
hf download nyu-visionx/Cambrian-S-3M --repo-type dataset --local-dir Cambrian-S-3M
# Decompress all archives
cd Cambrian-S-3M
bash decompress.sh
See decompress.sh for details.
cd $PATH_TO_LOCAL_DATASET
# Download LLaVA-Video-178K
hf download lmms-lab/LLaVA-Video-178K --repo-type dataset --local-dir LLaVA-Video-178K
# Decompress all archives
cd LLaVA-Video-178K
find . -name "*.tar.gz" -exec tar -zxf {} \;
cd $PATH_TO_LOCAL_DATASET
# Download LLaVA-Hound (ShareGPTVideo)
hf download ShareGPTVideo/train_video_and_instruction --repo-type dataset --local-dir train_video_and_instruction
# Decompress all archives
cd train_video_and_instruction/train_300k
find . -name "*.tar.gz" -exec tar -zxf {} \;
Link the additional datasets into the Cambrian-S-3M directory:
cd $PATH_TO_LOCAL_DATASET/Cambrian-S-3M
# Link LLaVA-Video-178K files
ln -s $PATH_TO_LOCAL_DATASET/LLaVA-Video-178K/* ./
# Link LLaVA-Hound frames
mkdir -p shareVideoGPTV/frames
ln -s $PATH_TO_LOCAL_DATASET/train_video_and_instruction/train_300k ./shareVideoGPTV/frames/all_frames
Run the sanity check to ensure all data is properly set up:
cd $PATH_TO_LOCAL_DATASET/Cambrian-S-3M
python sanity_check.py
hf command: Install with pip install -U "huggingface_hub[cli]"$PATH_TO_LOCAL_DATASETIf you use this dataset, please cite:
@article{yang2025cambrians,
title={Cambrian-S: Towards Spatial Supersensing in Video},
author={Yang, Shusheng and Yang, Jihan and Huang, Pinzhi and Brown, Ellis and Yang, Zihao and Yu, Yue and Tong, Shengbang and Zheng, Zihan and Xu, Yifan and Wang, Muhan and Lu, Daohan and Fergus, Rob and LeCun, Yann and Fei-Fei, Li and Xie, Saining},
journal={arXiv preprint arXiv:2511.04670},
year={2025}
}
This dataset is released under the Apache 2.0 license.
500 commits