Official codebase for Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors (Table1).
Python
0
23 commits
updated Sep 7, 2025
.cache/qwen2_5scripts/pool_full/finetune_poolbranch_3b_full_pool4.shscripts/eval_pool_full/eval_sub_baseline_qwen3b_pool4.sh$MVBench_{token ~ aware}$bash mvbench_hard.sh --pool2 <path_to_pool2_results_folder> --output <output_json_path>
The pool2 folder should contain multiple JSON files; the script merges them into one.pretrain folder.bash scripts/pretrain_projector_image_encoder.sh
bash scripts/pretrain_projector_video_encoder.sh
VideoGPT-plus
├─ .cache
│ ├─ instruction_data
│ ├─ pretraining
│ ├─ COCO
│ ├─ cc3M
pip install flash-attn --no-build-isolationpip install -r requirements.txtscripts/pretrain_projector_image_encoder.shscripts/pretrain_projector_video_encoder.shscripts/pool_fullexport PYTHONPATH="./VideoGPT-plus:$PYTHONPATH"
def divide_and_round(a, b):
result = a / b
if result < 1:
return 1
else:
return round(result)
video_feature: shape=(divide_and_round(16, pool_level), divide_and_round(16, pool_level))
image_feature: shape=(divide_and_round(24, pool_level), divide_and_round(24, pool_level))
pool1 (i.e., b=1) means no pooling is applied.scripts/eval_pool_fullPython
92.9%
Shell
7.1%
Official codebase for Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors (Table1).
Python
0
23 commits
updated Sep 7, 2025
.cache/qwen2_5scripts/pool_full/finetune_poolbranch_3b_full_pool4.shscripts/eval_pool_full/eval_sub_baseline_qwen3b_pool4.sh$MVBench_{token ~ aware}$bash mvbench_hard.sh --pool2 <path_to_pool2_results_folder> --output <output_json_path>
The pool2 folder should contain multiple JSON files; the script merges them into one.pretrain folder.bash scripts/pretrain_projector_image_encoder.sh
bash scripts/pretrain_projector_video_encoder.sh
VideoGPT-plus
├─ .cache
│ ├─ instruction_data
│ ├─ pretraining
│ ├─ COCO
│ ├─ cc3M
pip install flash-attn --no-build-isolationpip install -r requirements.txtscripts/pretrain_projector_image_encoder.shscripts/pretrain_projector_video_encoder.shscripts/pool_fullexport PYTHONPATH="./VideoGPT-plus:$PYTHONPATH"
def divide_and_round(a, b):
result = a / b
if result < 1:
return 1
else:
return round(result)
video_feature: shape=(divide_and_round(16, pool_level), divide_and_round(16, pool_level))
image_feature: shape=(divide_and_round(24, pool_level), divide_and_round(24, pool_level))
pool1 (i.e., b=1) means no pooling is applied.scripts/eval_pool_fullPython
92.9%
Shell
7.1%