The evaluation code code for UWLM dataset.
Jupyter Notebook
46
15 commits
updated May 8, 2026
π Kaggle Website β’ π€ Dataset on Hugging Face (on going) β’ π Paper
UVLM is a benchmark for underwater video-language understanding. It contains 2,109 videos, ~0.86M frames, 419 marine species/categories, and 20 fine-grained tasks covering both biological and environmental aspects of underwater scenes.
This repository provides:
uvlm-benchmark/
βββ README.md
βββ setup.py # optional, to install as a package
βββ uvlm_benchmark/
β βββ __init__.py
β βββ data/
β β βββ downloader.py # download from Hugging Face
β β βββ dataset.py # unified dataset loader
β β βββ schemas.py # annotation field definitions
β βββ eval/
β β βββ eval_mcqa.py
β β βββ eval_llm_judge.py
β β βββ metrics.py
β βββ utils/
β βββ io.py
βββ scripts/
β βββ download_from_hf.sh # one-click HF download
β βββ run_eval_mcqa.sh
βββ dataset/
β βββ README.md # explains real data layout on HF
β βββ annotations_example.json
βββ submissions/
β βββ sample_results.json
βββ docs/
β βββ index.md # GitHub Pages (optional)
βββ CITATION.cff
All large files (videos, full annotations, splits) are hosted on Hugging Face.
This repository only contains examples and format descriptions.
Option A: clone from HF directly
git lfs install
git clone https://huggingface.co/datasets/ZhouYang2002/UVLM
Option B: use our script (recommended)
bash scripts/download_from_hf.sh
After download, your data directory should look like:
uvlm/
βββ videos/ # optional: raw or segmented videos
βββ frames/ # optional: extracted frames
βββ annotations/ # all .json annotation files
βββ splits/ # train / val / test split files
Note:
dataset/annotations_example.jsonin this repo is a toy example to show the JSON schema.
Please download the real annotations from Hugging Face.
Install the package locally:
pip install -e .
This installs the package uvlm_benchmark, so you can run evaluation as modules:
python -m uvlm_benchmark.eval.eval_mcqa --help
UVLM is organized around underwater perception and reasoning. Each video may contain multiple QA items.
Example annotation (simplified):
{
"video_id": "video_00001",
"fps": 25,
"duration": 8.2,
"scene_type": "marine",
"qas": [
{
"qa_id": "video_00001_q1",
"question": "What species is shown in the video?",
"options": ["lionfish", "clownfish", "seahorse", "octopus"],
"answer": "clownfish",
"task_type": "species_identification",
"source": "human"
},
{
"qa_id": "video_00001_q2",
"question": "How clear is the water environment?",
"answer": "The water is slightly turbid.",
"task_type": "environment_description",
"source": "gpt4o"
}
]
}
Key fields
video_id: unique ID of the videoqas: list of QA items linked to the videotask_type: tells the evaluator which metric to use (MCQA vs LLM-based)source: where the QA comes from (human / GPT / expert)A full field-by-field description should be placed in dataset/README.md.
We provide two official evaluation routes to match the paper.
Use this for multiple-choice questions.
python -m uvlm_benchmark.eval.eval_mcqa --pred submissions/sample_results.json --gt /path/to/uvlm/annotations
--pred: your model predictions--gt: ground-truth annotations downloaded from HFPrediction file format (submissions/sample_results.json):
[
{
"video_id": "video_00001",
"qa_id": "video_00001_q1",
"pred_answer": "clownfish"
}
]
You can extend it with confidence scores if needed.
Some underwater questions require open-ended or knowledge-grounded evaluation.
python -m uvlm_benchmark.eval.eval_llm_judge --pred submissions/sample_results.json --gt /path/to/uvlm/annotations --model gpt-4o
uvlm_benchmark/eval/prompts/.We start with a simple Markdown leaderboard in this repo.
| Rank | Method | MCA | FGC | Overall | Link |
|---|---|---|---|---|---|
| 1 | GPT-4o (official) | 77.72 | 81.47 | 77.95 | - |
| 2 | VideoLLaMA3-7B + UVLM | 76.85 | 57.17 | 73.04 | code |
| 3 | InternVL2.5-8B + UVLM | 70.26 | 43.94 | 69.45 | code |
How to submit
results.jsonLater, this table can be moved to GitHub Pages / HF Spaces.
We provide a tiny helper to download from HF inside the codebase:
from uvlm_benchmark.data.downloader import download_uvlm
download_uvlm(
dest="./data/uvlm",
repo_id="your-org/uvlm",
)
Example implementation (uvlm_benchmark/data/downloader.py):
from huggingface_hub import snapshot_download
from pathlib import Path
def download_uvlm(dest="./data/uvlm", repo_id="your-org/uvlm", revision=None):
dest = Path(dest)
dest.mkdir(parents=True, exist_ok=True)
snapshot_download(
repo_id=repo_id,
local_dir=str(dest),
repo_type="dataset",
revision=revision,
local_dir_use_symlinks=False,
)
return dest
This way, real data stays on HF while GitHub only stores code and examples.
If you want a nicer βofficialβ page:
docs/index.md.main β /docs.https://yourname.github.io/uvlm-benchmark/) will become your benchmark homepage.If you use UVLM in your research, please cite both the conference version and the arXiv version.
@article{xue2026uvlm,
title = {UVLM: Benchmarking Video-Language Model for Underwater World Understanding},
author = {Xizhe Xue and Yangzhou and Dawei Yan and Lijie Tao and Junjie Li and Ying Li and Haokui Zhang and Rong Xiao},
journal = {Proceedings of the AAAI Conference on Artificial Intelligence},
year = {2026}
}
The evaluation code code for UWLM dataset.
Jupyter Notebook
46
15 commits
updated May 8, 2026
π Kaggle Website β’ π€ Dataset on Hugging Face (on going) β’ π Paper
UVLM is a benchmark for underwater video-language understanding. It contains 2,109 videos, ~0.86M frames, 419 marine species/categories, and 20 fine-grained tasks covering both biological and environmental aspects of underwater scenes.
This repository provides:
uvlm-benchmark/
βββ README.md
βββ setup.py # optional, to install as a package
βββ uvlm_benchmark/
β βββ __init__.py
β βββ data/
β β βββ downloader.py # download from Hugging Face
β β βββ dataset.py # unified dataset loader
β β βββ schemas.py # annotation field definitions
β βββ eval/
β β βββ eval_mcqa.py
β β βββ eval_llm_judge.py
β β βββ metrics.py
β βββ utils/
β βββ io.py
βββ scripts/
β βββ download_from_hf.sh # one-click HF download
β βββ run_eval_mcqa.sh
βββ dataset/
β βββ README.md # explains real data layout on HF
β βββ annotations_example.json
βββ submissions/
β βββ sample_results.json
βββ docs/
β βββ index.md # GitHub Pages (optional)
βββ CITATION.cff
All large files (videos, full annotations, splits) are hosted on Hugging Face.
This repository only contains examples and format descriptions.
Option A: clone from HF directly
git lfs install
git clone https://huggingface.co/datasets/ZhouYang2002/UVLM
Option B: use our script (recommended)
bash scripts/download_from_hf.sh
After download, your data directory should look like:
uvlm/
βββ videos/ # optional: raw or segmented videos
βββ frames/ # optional: extracted frames
βββ annotations/ # all .json annotation files
βββ splits/ # train / val / test split files
Note:
dataset/annotations_example.jsonin this repo is a toy example to show the JSON schema.
Please download the real annotations from Hugging Face.
Install the package locally:
pip install -e .
This installs the package uvlm_benchmark, so you can run evaluation as modules:
python -m uvlm_benchmark.eval.eval_mcqa --help
UVLM is organized around underwater perception and reasoning. Each video may contain multiple QA items.
Example annotation (simplified):
{
"video_id": "video_00001",
"fps": 25,
"duration": 8.2,
"scene_type": "marine",
"qas": [
{
"qa_id": "video_00001_q1",
"question": "What species is shown in the video?",
"options": ["lionfish", "clownfish", "seahorse", "octopus"],
"answer": "clownfish",
"task_type": "species_identification",
"source": "human"
},
{
"qa_id": "video_00001_q2",
"question": "How clear is the water environment?",
"answer": "The water is slightly turbid.",
"task_type": "environment_description",
"source": "gpt4o"
}
]
}
Key fields
video_id: unique ID of the videoqas: list of QA items linked to the videotask_type: tells the evaluator which metric to use (MCQA vs LLM-based)source: where the QA comes from (human / GPT / expert)A full field-by-field description should be placed in dataset/README.md.
We provide two official evaluation routes to match the paper.
Use this for multiple-choice questions.
python -m uvlm_benchmark.eval.eval_mcqa --pred submissions/sample_results.json --gt /path/to/uvlm/annotations
--pred: your model predictions--gt: ground-truth annotations downloaded from HFPrediction file format (submissions/sample_results.json):
[
{
"video_id": "video_00001",
"qa_id": "video_00001_q1",
"pred_answer": "clownfish"
}
]
You can extend it with confidence scores if needed.
Some underwater questions require open-ended or knowledge-grounded evaluation.
python -m uvlm_benchmark.eval.eval_llm_judge --pred submissions/sample_results.json --gt /path/to/uvlm/annotations --model gpt-4o
uvlm_benchmark/eval/prompts/.We start with a simple Markdown leaderboard in this repo.
| Rank | Method | MCA | FGC | Overall | Link |
|---|---|---|---|---|---|
| 1 | GPT-4o (official) | 77.72 | 81.47 | 77.95 | - |
| 2 | VideoLLaMA3-7B + UVLM | 76.85 | 57.17 | 73.04 | code |
| 3 | InternVL2.5-8B + UVLM | 70.26 | 43.94 | 69.45 | code |
How to submit
results.jsonLater, this table can be moved to GitHub Pages / HF Spaces.
We provide a tiny helper to download from HF inside the codebase:
from uvlm_benchmark.data.downloader import download_uvlm
download_uvlm(
dest="./data/uvlm",
repo_id="your-org/uvlm",
)
Example implementation (uvlm_benchmark/data/downloader.py):
from huggingface_hub import snapshot_download
from pathlib import Path
def download_uvlm(dest="./data/uvlm", repo_id="your-org/uvlm", revision=None):
dest = Path(dest)
dest.mkdir(parents=True, exist_ok=True)
snapshot_download(
repo_id=repo_id,
local_dir=str(dest),
repo_type="dataset",
revision=revision,
local_dir_use_symlinks=False,
)
return dest
This way, real data stays on HF while GitHub only stores code and examples.
If you want a nicer βofficialβ page:
docs/index.md.main β /docs.https://yourname.github.io/uvlm-benchmark/) will become your benchmark homepage.If you use UVLM in your research, please cite both the conference version and the arXiv version.
@article{xue2026uvlm,
title = {UVLM: Benchmarking Video-Language Model for Underwater World Understanding},
author = {Xizhe Xue and Yangzhou and Dawei Yan and Lijie Tao and Junjie Li and Ying Li and Haokui Zhang and Rong Xiao},
journal = {Proceedings of the AAAI Conference on Artificial Intelligence},
year = {2026}
}