Model type: VTSUM-BLIP is an end-to-end cross-modal video summarization model.
Model description:
The file structure of Model zoo looks like:
outputs
├── blip
│ └── model_base_capfilt_large.pth
├── vt_clipscore
│ └── vt_clip.pth
├── vtsum_tt
│ └── vtsum_tt.pth
└── vtsum_tt_ca
└── vtsum_tt_ca.pth
Paper or resources for more information: https://videoxum.github.io/
4 commits
Model type: VTSUM-BLIP is an end-to-end cross-modal video summarization model.
Model description:
The file structure of Model zoo looks like:
outputs
├── blip
│ └── model_base_capfilt_large.pth
├── vt_clipscore
│ └── vt_clip.pth
├── vtsum_tt
│ └── vtsum_tt.pth
└── vtsum_tt_ca
└── vtsum_tt_ca.pth
Paper or resources for more information: https://videoxum.github.io/
4 commits