Multimodal Benchmarks & Video Understanding

2 repos

This cluster focuses on benchmarking and evaluation frameworks for multimodal AI systems, particularly those involving video understanding, temporal action localization, and action recognition. The repos center on creating standardized datasets and evaluation protocols for assessing model performance across vision, language, and audio modalities. Researchers exploring this area will find comprehensive benchmark suites, evaluation toolkits, and dataset resources for validating multimodal models in real-world scenarios.

HTML · 1
benchmark ·7
document-ai ·7
document-extraction ·7
form-understanding ·7
json-schema ·7
multimodal ·7
structured-extraction ·7

Related clusters