18 repos
Benchmark suites and evaluation frameworks for assessing the performance and capabilities of large language models and vision-language models across diverse domains. These repositories provide standardized evaluation datasets, metrics, and testing protocols—spanning medical imaging, code generation, multimodal understanding, and cyber security—enabling systematic comparison of model behavior and safety properties. Researchers and practitioners use these benchmarks to measure model improvements, identify failure modes, and validate alignment with real-world requirements.