[NeurIPS 2026] V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
See the code1 Sun Yat-sen University Shenzhen Campus 2 Shenzhen Loop Area Institute
3 Sichuan University 4 Shanghai Jiao Tong University
⚡ A training-free and plug-and-play Curvature-Aware Spatio-Temporal pruning framework for efficient long-context video inference.
2026.09.25 🎉🎉 Our V-CAST has been accepted to NeurIPS 2026!2026.03.26 🤗🤗 We opened the V-CAST repository.
V-CAST is a training-free and plug-and-play curvature-aware spatio-temporal pruning framework for efficient long-context video inference. It revisits token compression from the perspective of spatio-temporal information coverage, and combines:
(t, h, w) grid.
git clone https://github.com/xinyouu/V-CAST.git
cd V-CAST
conda create -n vcast python=3.10 -y
conda activate vcast
pip install --upgrade pip
pip install -e ".[train]"
If you want to measure the latency and GPU memory, please use the custom installation.
cd lmms-eval
pip install -e .
Or you can also use the official installation.
pip install git+https://github.com/EvolvingLMMs-Lab/lmms-eval.git
| Model Base | Status | Code Path |
|---|---|---|
| Qwen3-VL | ✅ Released | compressor/v_cast/modeling_qwen3_vl_v_cast.py |
| LLaVA-OneVision / LLaVA-Video | 🚧 Planned | - |
| Qwen2.5-Omni / Qwen3-Omni | 🚧 Planned | - |
Our evaluation pipeline is built on top of lmms-eval, a unified toolkit for multimodal evaluation across text, image, video, and audio tasks. The current public release focuses on the Qwen3-VL evaluation path with V-CAST enabled by default, while additional model-family integrations will be released in follow-up updates. Some comparison baselines are also available in the open-source VidCom2 repository.
bash examples/v_cast/inference_qwen3vl_v_cast_64.sh
| Component | Path |
|---|---|
| V-CAST wrapper | compressor/v_cast/main.py |
| V-CAST core implementation | compressor/v_cast/modeling_qwen3_vl_v_cast.py |
| Qwen3-VL evaluation wrapper | lmms_eval/models/simple/qwen3_vl.py |
| Example script | examples/v_cast/inference_qwen3vl_v_cast_64.sh |
If you find this repository useful, please cite:
@article{lin2026vcast,
title={V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models},
author={Lin, Xinying and Liu, Xuyang and Wang, Yiyu and Ma, Teng and Ren, Wenqi},
journal={arXiv preprint arXiv:2603.27650},
year={2026}
}
We extend our gratitude to the open-source efforts of LLaVA-OneVision and Qwen3-VL.
For any question about our paper or code, please email xinyinglin@slai.edu.cn.
Python
99.9%
[NeurIPS 2026] V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
See the code1 Sun Yat-sen University Shenzhen Campus 2 Shenzhen Loop Area Institute
3 Sichuan University 4 Shanghai Jiao Tong University
⚡ A training-free and plug-and-play Curvature-Aware Spatio-Temporal pruning framework for efficient long-context video inference.
2026.09.25 🎉🎉 Our V-CAST has been accepted to NeurIPS 2026!2026.03.26 🤗🤗 We opened the V-CAST repository.
V-CAST is a training-free and plug-and-play curvature-aware spatio-temporal pruning framework for efficient long-context video inference. It revisits token compression from the perspective of spatio-temporal information coverage, and combines:
(t, h, w) grid.
git clone https://github.com/xinyouu/V-CAST.git
cd V-CAST
conda create -n vcast python=3.10 -y
conda activate vcast
pip install --upgrade pip
pip install -e ".[train]"
If you want to measure the latency and GPU memory, please use the custom installation.
cd lmms-eval
pip install -e .
Or you can also use the official installation.
pip install git+https://github.com/EvolvingLMMs-Lab/lmms-eval.git
| Model Base | Status | Code Path |
|---|---|---|
| Qwen3-VL | ✅ Released | compressor/v_cast/modeling_qwen3_vl_v_cast.py |
| LLaVA-OneVision / LLaVA-Video | 🚧 Planned | - |
| Qwen2.5-Omni / Qwen3-Omni | 🚧 Planned | - |
Our evaluation pipeline is built on top of lmms-eval, a unified toolkit for multimodal evaluation across text, image, video, and audio tasks. The current public release focuses on the Qwen3-VL evaluation path with V-CAST enabled by default, while additional model-family integrations will be released in follow-up updates. Some comparison baselines are also available in the open-source VidCom2 repository.
bash examples/v_cast/inference_qwen3vl_v_cast_64.sh
| Component | Path |
|---|---|
| V-CAST wrapper | compressor/v_cast/main.py |
| V-CAST core implementation | compressor/v_cast/modeling_qwen3_vl_v_cast.py |
| Qwen3-VL evaluation wrapper | lmms_eval/models/simple/qwen3_vl.py |
| Example script | examples/v_cast/inference_qwen3vl_v_cast_64.sh |
If you find this repository useful, please cite:
@article{lin2026vcast,
title={V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models},
author={Lin, Xinying and Liu, Xuyang and Wang, Yiyu and Ma, Teng and Ren, Wenqi},
journal={arXiv preprint arXiv:2603.27650},
year={2026}
}
We extend our gratitude to the open-source efforts of LLaVA-OneVision and Qwen3-VL.
For any question about our paper or code, please email xinyinglin@slai.edu.cn.
Python
99.9%