RouteWorks/RouterArena

Dataset

Bloom Taxonomy Filtered Dataset

9

12 commits

1 linked in READMEs

updated Dec 3, 2025

See the code

README

Bloom Taxonomy Filtered Dataset

Paper | Project Page | Code

Description

This is a testing dataset for LLM routers. It contains 8.4k questions from 9 domains, and 44 categories (defined by Dewey Decimal Classes).

Dataset Overview

This dataset contains 8,400 carefully selected questions from the RouterEvalBenchmark, organized by:

  • Subject Categories: Computer Science, Philosophy, Social Science, Language, Science, Technology, Arts, Literature, History
  • Bloom Taxonomy Levels: Remember, Understand, Apply, Analyze, Evaluate, Create
  • Cosine Similarity Deduplication: For most datasets, questions are selected based on cosine similarity to ensure diversity
  • Random Selection: For specific datasets (ChessInstruct, SuperGLUE*, WMT19*), random selection is used

Usage

from datasets import load_dataset
dataset = load_dataset("RouteWorks/RouterArena")

GitHub:

Detailed usage could be found on our GitHub.

Citation

If you find our work useful, please cite:

@misc{lu2025routerarenaopenplatformcomprehensive,
  title        = {RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers},
  author       = {Yifan Lu and Rixin Liu and Jiayi Yuan and Xingqi Cui and Shenrun Zhang and Hongyi Liu and Jiarong Xing},
  year         = {2025},
  eprint       = {2510.00202},
  archivePrefix= {arXiv},
  primaryClass = {cs.LG},
  url          = {https://arxiv.org/abs/2510.00202}
}

RouteWorks/RouterArena

Dataset

Bloom Taxonomy Filtered Dataset

9

12 commits

1 linked in READMEs

updated Dec 3, 2025

See the code

README

Bloom Taxonomy Filtered Dataset

Paper | Project Page | Code

Description

This is a testing dataset for LLM routers. It contains 8.4k questions from 9 domains, and 44 categories (defined by Dewey Decimal Classes).

Dataset Overview

This dataset contains 8,400 carefully selected questions from the RouterEvalBenchmark, organized by:

  • Subject Categories: Computer Science, Philosophy, Social Science, Language, Science, Technology, Arts, Literature, History
  • Bloom Taxonomy Levels: Remember, Understand, Apply, Analyze, Evaluate, Create
  • Cosine Similarity Deduplication: For most datasets, questions are selected based on cosine similarity to ensure diversity
  • Random Selection: For specific datasets (ChessInstruct, SuperGLUE*, WMT19*), random selection is used

Usage

from datasets import load_dataset
dataset = load_dataset("RouteWorks/RouterArena")

GitHub:

Detailed usage could be found on our GitHub.

Citation

If you find our work useful, please cite:

@misc{lu2025routerarenaopenplatformcomprehensive,
  title        = {RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers},
  author       = {Yifan Lu and Rixin Liu and Jiayi Yuan and Xingqi Cui and Shenrun Zhang and Hongyi Liu and Jiarong Xing},
  year         = {2025},
  eprint       = {2510.00202},
  archivePrefix= {arXiv},
  primaryClass = {cs.LG},
  url          = {https://arxiv.org/abs/2510.00202}
}