🏆 🥇 Winner Solution for ICCV VQualA 2025 EVQA-SnapUGC Challenge at VQualA 2025 Workshop @ ICCV 2025
Python
20
5 commits
updated Aug 8, 2025
🏆 🥇 Winner Solution for ICCV VQualA 2025 EVQA-SnapUGC Challenge at VQualA 2025 Workshop @ ICCV 2025
Official Implementation for Engagement Prediction of Short Videos with Large Multimodal Models
The rapid proliferation of user-generated content (UGC) on short-form video platforms has made video engagement prediction increasingly important for optimizing recommendation systems and guiding content creation. However, this task remains challenging due to the complex interplay of factors such as semantic content, visual quality, audio characteristics, and user background.
In this work, we empirically investigate the potential of large multimodal models (LMMs) for video engagement prediction. We adopt two representative LMMs:
Both models demonstrate competitive performance against state-of-the-art baselines, showcasing the effectiveness of LMMs in engagement prediction. Notably, VideoLLaMA2 consistently outperforms Qwen2.5-VL, highlighting the importance of audio features in engagement prediction.
| Team Name | Final Score | SROCC | PLCC |
|---|---|---|---|
| 🥇 ECNU-SJTU VQA (Ours) | 0.710 | 0.707 | 0.714 |
| IMCL-DAMO | 0.698 | 0.696 | 0.702 |
| HKUST-Cardiff-MI-BAAI | 0.680 | 0.677 | 0.684 |
| MCCE | 0.667 | 0.660 | 0.668 |
| EasyVQA | 0.667 | 0.664 | 0.671 |
| Rochester | 0.449 | 0.405 | 0.515 |
| brucelyu | 0.441 | 0.439 | 0.444 |
Note: Final score = 0.6 × SROCC + 0.4 × PLCC
For more detailed results, please refer to the challenge report paper.
edts)28tq)SnapUGC/
├── train_data.csv # Training annotations
├── val_data.csv # Validation annotations
├── train_videos/ # Training video files
└── val_videos/ # Validation video files
cd VideoLLaMA2-audio_visual
# Create conda environment
conda create -n videollama2 python=3.9
conda activate videollama2
# Install dependencies
pip install -r requirements.txt
# Download model weights
python download_model_weight.py
# --local_dir specifies the target directory for model weights
# Prepare training dataset
python prepare_dataset.py
# VIDEO_DIR = " " # Path to video directory
# AUDIO_DIR = " " # Audio directory parameter (deprecated) - VideoLLaMA2 automatically extracts audio features from video files
# CSV_PATH = "train_data.csv" # Input CSV file containing video metadata
# OUTPUT_TRAIN = "train.json" # Output JSON file for training data
# Start training
sh train.sh
# --model_path # Path to pre-trained model weights
# --data_path # Path to training data JSON file
# --num_frames # Number of frames to extract from each video (first N frames)
# --output_dir # Directory for saving trained model weights
# --num_train_epochs # Total number of training epochs
# --per_device_train_batch_size # Batch size per GPU device
Batch Testing:
# Generate validation dataset
python prepare_dataset.py
# Run validation
sh run_validation.sh
# --model-path # Path to the trained model weights
# --modal-type # Specify "av" mode for audio-visual evaluation
# --json_file # Path to the validation JSON file
Single Video Testing:
python videollama2/test_single_video.py \
--model-path /path/to/trained/model \
--modal-type av \
--video-path /path/to/video.mp4 \
--title "Your Video Title" \
--description "Your Video Description"
Parameters:
--model-path: Path to trained model weights (required)--modal-type: Modal type - "a" (audio), "v" (video), "av" (audio-visual)--video-path: Path to video file (required)--title: Video title (optional, default: None)--description: Video description (optional, default: None)cd Qwen2.5-VL
# Create conda environment
conda create -n qwenvl python=3.9
conda activate qwenvl
# Follow official Qwen2.5-VL installation
# https://github.com/QwenLM/Qwen2.5-VL
# Prepare dataset for testing
python prepare_dataset_qwenvl.py
# VIDEO_DIR = " " # Path to the folder containing video files
# CSV_PATH = " " # Path to the provided CSV file
# OUTPUT_TRAIN = " " # Path to the output JSON file used for testing
Batch Testing:
python infer_evqa.py \
--model-path /path/to/trained/model \
--video-folder /path/to/video/folder \
--question-file /path/to/question.json \
--save-csv /path/to/result.csv
Single Video Testing:
python test_single_video_qwenvl.py \
--model-path /path/to/trained/model \
--video-path /path/to/video.mp4 \
--title "Your Video Title" \
--description "Your Video Description"
Parameters:
--model-path: Path to trained Qwen2.5-VL model weights (required)--video-path: Path to video file (required)--title: Video title (optional, default: None)--description: Video description (optional, default: None)3aqc)98zr)If you find this code useful for your research, please cite our paper:
@inproceedings{sun2025engagement,
title={Engagement Prediction of Short Videos with Large Multimodal Models},
author={Sun, Wei and Cao, Linhan and Cao, Yuqin and Zhang, Weixia and Wen, Wen and Zhang, Kaiwei and Chen, Zijian and Lu, Fangfang and Min, Xiongkuo and Zhai, Guangtao},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision (ICCV) Workshops},
pages={1--10},
year={2025}
}
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
⭐ Star this repository if you find it helpful!
🏆 🥇 Winner Solution for ICCV VQualA 2025 EVQA-SnapUGC Challenge at VQualA 2025 Workshop @ ICCV 2025
Python
20
5 commits
updated Aug 8, 2025
🏆 🥇 Winner Solution for ICCV VQualA 2025 EVQA-SnapUGC Challenge at VQualA 2025 Workshop @ ICCV 2025
Official Implementation for Engagement Prediction of Short Videos with Large Multimodal Models
The rapid proliferation of user-generated content (UGC) on short-form video platforms has made video engagement prediction increasingly important for optimizing recommendation systems and guiding content creation. However, this task remains challenging due to the complex interplay of factors such as semantic content, visual quality, audio characteristics, and user background.
In this work, we empirically investigate the potential of large multimodal models (LMMs) for video engagement prediction. We adopt two representative LMMs:
Both models demonstrate competitive performance against state-of-the-art baselines, showcasing the effectiveness of LMMs in engagement prediction. Notably, VideoLLaMA2 consistently outperforms Qwen2.5-VL, highlighting the importance of audio features in engagement prediction.
| Team Name | Final Score | SROCC | PLCC |
|---|---|---|---|
| 🥇 ECNU-SJTU VQA (Ours) | 0.710 | 0.707 | 0.714 |
| IMCL-DAMO | 0.698 | 0.696 | 0.702 |
| HKUST-Cardiff-MI-BAAI | 0.680 | 0.677 | 0.684 |
| MCCE | 0.667 | 0.660 | 0.668 |
| EasyVQA | 0.667 | 0.664 | 0.671 |
| Rochester | 0.449 | 0.405 | 0.515 |
| brucelyu | 0.441 | 0.439 | 0.444 |
Note: Final score = 0.6 × SROCC + 0.4 × PLCC
For more detailed results, please refer to the challenge report paper.
edts)28tq)SnapUGC/
├── train_data.csv # Training annotations
├── val_data.csv # Validation annotations
├── train_videos/ # Training video files
└── val_videos/ # Validation video files
cd VideoLLaMA2-audio_visual
# Create conda environment
conda create -n videollama2 python=3.9
conda activate videollama2
# Install dependencies
pip install -r requirements.txt
# Download model weights
python download_model_weight.py
# --local_dir specifies the target directory for model weights
# Prepare training dataset
python prepare_dataset.py
# VIDEO_DIR = " " # Path to video directory
# AUDIO_DIR = " " # Audio directory parameter (deprecated) - VideoLLaMA2 automatically extracts audio features from video files
# CSV_PATH = "train_data.csv" # Input CSV file containing video metadata
# OUTPUT_TRAIN = "train.json" # Output JSON file for training data
# Start training
sh train.sh
# --model_path # Path to pre-trained model weights
# --data_path # Path to training data JSON file
# --num_frames # Number of frames to extract from each video (first N frames)
# --output_dir # Directory for saving trained model weights
# --num_train_epochs # Total number of training epochs
# --per_device_train_batch_size # Batch size per GPU device
Batch Testing:
# Generate validation dataset
python prepare_dataset.py
# Run validation
sh run_validation.sh
# --model-path # Path to the trained model weights
# --modal-type # Specify "av" mode for audio-visual evaluation
# --json_file # Path to the validation JSON file
Single Video Testing:
python videollama2/test_single_video.py \
--model-path /path/to/trained/model \
--modal-type av \
--video-path /path/to/video.mp4 \
--title "Your Video Title" \
--description "Your Video Description"
Parameters:
--model-path: Path to trained model weights (required)--modal-type: Modal type - "a" (audio), "v" (video), "av" (audio-visual)--video-path: Path to video file (required)--title: Video title (optional, default: None)--description: Video description (optional, default: None)cd Qwen2.5-VL
# Create conda environment
conda create -n qwenvl python=3.9
conda activate qwenvl
# Follow official Qwen2.5-VL installation
# https://github.com/QwenLM/Qwen2.5-VL
# Prepare dataset for testing
python prepare_dataset_qwenvl.py
# VIDEO_DIR = " " # Path to the folder containing video files
# CSV_PATH = " " # Path to the provided CSV file
# OUTPUT_TRAIN = " " # Path to the output JSON file used for testing
Batch Testing:
python infer_evqa.py \
--model-path /path/to/trained/model \
--video-folder /path/to/video/folder \
--question-file /path/to/question.json \
--save-csv /path/to/result.csv
Single Video Testing:
python test_single_video_qwenvl.py \
--model-path /path/to/trained/model \
--video-path /path/to/video.mp4 \
--title "Your Video Title" \
--description "Your Video Description"
Parameters:
--model-path: Path to trained Qwen2.5-VL model weights (required)--video-path: Path to video file (required)--title: Video title (optional, default: None)--description: Video description (optional, default: None)3aqc)98zr)If you find this code useful for your research, please cite our paper:
@inproceedings{sun2025engagement,
title={Engagement Prediction of Short Videos with Large Multimodal Models},
author={Sun, Wei and Cao, Linhan and Cao, Yuqin and Zhang, Weixia and Wen, Wen and Zhang, Kaiwei and Chen, Zijian and Lu, Fangfang and Min, Xiongkuo and Zhai, Guangtao},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision (ICCV) Workshops},
pages={1--10},
year={2025}
}
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
⭐ Star this repository if you find it helpful!