lucas-ventura/chapter-llama

Dataset

1

stars

500

commits

1

linked in READMEs

Oct 20, 2025

updated

chaptering
VidChapters
video
video-chaptering

README

VidChapters Dataset for Chapter-Llama

This repository contains the dataset used in the paper "Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs" (CVPR 2025).

Overview

VidChapters-7M is a large-scale dataset for video chaptering, containing:

  • ~817k videos with ASR data (~20GB)
  • Captions extracted from videos using various sampling strategies
  • Chapter annotations with timestamps and titles

Data Structure

The dataset is organized as follows:

  • ASR data: Speech transcripts with timestamps
  • Chapter data: Chapter boundaries and titles
  • Captions: Visual frame captions captured at strategic timestamps
  • Subsets: Various pre-defined subsets for training/validation/testing

Usage

This dataset is designed to be used with the Chapter-Llama codebase, which provides tools for:

  • Loading and processing the dataset
  • Training LLM-based chaptering models
  • Evaluating chaptering performance

Citation

If you use this dataset in your work, please cite our paper:

@article{ventura25chapter,
    title     = {{Chapter-Llama}: Efficient Chaptering in Hour-Long Videos with {LLM}s},
    author    = {Lucas Ventura and Antoine Yang and Cordelia Schmid and G{\"u}l Varol},
    journal   = {CVPR},
    year      = {2025}
}

License

This dataset is distributed under an MIT License. Please check the repository for more details.

Contributors

lucas-ventura

500 commits

lucas-ventura/chapter-llama

Dataset

1

stars

500

commits

1

linked in READMEs

Oct 20, 2025

updated

chaptering
VidChapters
video
video-chaptering

README

VidChapters Dataset for Chapter-Llama

This repository contains the dataset used in the paper "Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs" (CVPR 2025).

Overview

VidChapters-7M is a large-scale dataset for video chaptering, containing:

  • ~817k videos with ASR data (~20GB)
  • Captions extracted from videos using various sampling strategies
  • Chapter annotations with timestamps and titles

Data Structure

The dataset is organized as follows:

  • ASR data: Speech transcripts with timestamps
  • Chapter data: Chapter boundaries and titles
  • Captions: Visual frame captions captured at strategic timestamps
  • Subsets: Various pre-defined subsets for training/validation/testing

Usage

This dataset is designed to be used with the Chapter-Llama codebase, which provides tools for:

  • Loading and processing the dataset
  • Training LLM-based chaptering models
  • Evaluating chaptering performance

Citation

If you use this dataset in your work, please cite our paper:

@article{ventura25chapter,
    title     = {{Chapter-Llama}: Efficient Chaptering in Hour-Long Videos with {LLM}s},
    author    = {Lucas Ventura and Antoine Yang and Cordelia Schmid and G{\"u}l Varol},
    journal   = {CVPR},
    year      = {2025}
}

License

This dataset is distributed under an MIT License. Please check the repository for more details.

Contributors

lucas-ventura

500 commits