14 repos
Benchmark leaderboards and evaluation frameworks for assessing AI models across diverse modalities and tasks. These repositories provide standardized benchmarking infrastructure using Gradio for interactive web interfaces, enabling researchers to compare model performance on tasks like image generation, video synthesis, visual reasoning, and multimodal understanding. The cluster consolidates leaderboard platforms and associated evaluation suites that have become central to tracking progress in generative AI.