RouterArena: An open framework for evaluating LLM routers with standardized datasets, metrics, an automated framework, and a live leaderboard.
See the code
RouterArena is an open evaluation platform and leaderboard for LLM routers—systems that automatically select the best model for a given query. As the LLM ecosystem diversifies with models varying in size, capability, and cost, routing has become critical for balancing performance and cost. Yet, LLM routers currently lack a standardized evaluation framework to assess how effectively they trade off accuracy, cost, and other related metrics.
RouterArena bridges this gap by providing an open evaluation platform and benchmarking framework for both open-source and commercial routers. It has the following key features:
We aim for RouterArena to serve as a foundation for the community to evaluate, understand, and advance LLM routing systems.
[!IMPORTANT] RouterArena is an evaluation-only dataset. Submissions that train, fit, or tune any router component on RouterArena data (including the label files) will be rejected, and any accepted submission found in violation will be withdrawn.
For more details, please see our website and blog.
| Rank | Router | Affiliation | Acc-Cost Arena | Accuracy | Cost/1K Queries | Optimal Selection | Optimal Cost | Optimal Accuracy | Latency | Robustness |
|---|---|---|---|---|---|---|---|---|---|---|
| 🥇 | KT-ModelRouter | 👤 @DusanBaek | 76.28 | 78.14 | $0.27 | — | — | — | — | 80.48 |
| 🥈 | Sqwish Router | 👤 @namitha-sqwish | 76.21 | 79.76 | $0.70 | 9.04 | 23.49 | 94.07 | — | 51.67 |
| 🥉 | Divyam | 👤 @samikd | 75.85 | 78.59 | $0.48 | 4.93 | 16.03 | 93.43 | — | 98.33 |
| 4 | Cross-Router | 👤 @JiaHg | 75.75 | 78.14 | $0.40 | 17.66 | 45.49 | 90.31 | — | 67.14 |
| 5 | LLM Router [PyPI] | 👤 @ypollak2 | 75.69 | 78.44 | $0.49 | — | — | — | — | 100.00 |
| 6 | R**2: Reasonometry Router | 👤 @sashakolpakov | 75.53 | 78.11 | $0.45 | — | — | — | — | 93.57 |
| 7 | SapientAI Auto Router | 💼 Publicis Sapient | 75.41 | 77.89 | $0.43 | 5.93 | 19.99 | 92.00 | — | 73.33 |
| 8 | vLLM‑SR [Code] [HF] | 🎓 vLLM SR Team | 74.86 | 77.18 | $0.42 | 16.81 | 25.10 | 89.37 | — | 67.62 |
| 9 | nadir-caliper | 👤 @doramirdor | 74.55 | 75.84 | $0.22 | — | — | — | — | 79.76 |
| 10 | AgentForge Router | 👤 @YangY-Z | 74.13 | 74.72 | $0.13 | 17.84 | 52.47 | 98.68 | — | 40.48 |
| 11 | BARouter | 👤 @zulk2002 | 73.79 | 75.72 | $0.36 | 64.41 | 67.13 | 93.84 | — | 68.81 |
| 12 | Weave Router | 🎓 Weave | 72.82 | 76.32 | $0.94 | — | — | — | — | 100.00 |
| 13 | Nadir Router | 🎓 NadirRouter | 72.29 | 75.01 | $0.68 | — | — | — | — | 25.48 |
| 14 | OrcaRouter‑Adaptive [Code] [Paper] [X] | 🎓 Continuum AI | 72.08 | 75.54 | $1.00 | — | — | — | — | 22.62 |
| 15 | Hybrid Router | 👤 @mikemao27 | 72.08 | 71.38 | $0.04 | 89.87 | 94.19 | 92.81 | — | 96.67 |
| 16 | R2-Router | 🎓 UCF | 71.60 | 71.23 | $0.06 | 24.51 | 48.70 | 99.85 | — | 45.71 |
| 17 | cruq-router | 👤 @nabaruns | 70.77 | 71.35 | $0.18 | — | — | — | — | 81.67 |
| 18 | chuzom-solo-v32 | 👤 @ypollak2 | 70.61 | 70.59 | $0.10 | — | — | — | — | 100.00 |
| 19 | Azure-Model-Router [Web] | 💼 Microsoft | 70.42 | 72.94 | $0.73 | — | — | — | — | 71.43 |
| 20 | Auto Router | 👤 @cxf2015 | 70.05 | 70.17 | $0.12 | 37.58 | 40.02 | 86.04 | — | 49.52 |
| 21 | Lynkr | 👤 @vishalveerareddy123 | 67.65 | 68.41 | $0.29 | 10.97 | 16.08 | 84.48 | — | 92.38 |
| 22 | MIRT‑BERT [Code] | 🎓 USTC | 66.89 | 66.88 | $0.15 | 3.44 | 19.62 | 78.18 | 27.03 | 61.19 |
| 23 | NIRT‑BERT [Code] | 🎓 USTC | 66.12 | 66.34 | $0.21 | 3.83 | 14.04 | 77.88 | 10.42 | 49.29 |
| 24 | AsiaInfo-Router | 👤 @Uncle-LL | 65.87 | 75.20 | $8.54 | — | — | — | — | 69.52 |
| 25 | GPT‑5 | 💼 OpenAI | 64.32 | 73.96 | $10.02 | — | — | — | — | — |
| 26 | CARROT [Code] [HF] | 🎓 UMich | 63.87 | 67.21 | $2.06 | 2.68 | 6.77 | 78.63 | 1.50 | 89.05 |
| 27 | Chayan [HF] | 🎓 Adaptive Classifier | 63.83 | 64.89 | $0.56 | 43.03 | 43.75 | 88.74 | — | — |
| 28 | RouterBench‑MLP [Code] [HF] | 🎓 Martian | 57.56 | 61.62 | $4.83 | 13.39 | 24.45 | 83.32 | 90.91 | 80.00 |
| 29 | NotDiamond | 💼 NotDiamond | 57.29 | 60.83 | $4.10 | 1.55 | 2.14 | 76.81 | — | 55.91 |
| 30 | GraphRouter [Code] | 🎓 UIUC | 57.22 | 57.00 | $0.34 | 4.73 | 38.33 | 74.25 | 2.70 | 94.29 |
| 31 | RouterBench‑KNN [Code] [HF] | 🎓 Martian | 55.48 | 58.69 | $4.27 | 13.09 | 25.49 | 78.77 | 1.33 | 83.33 |
| 32 | RouteLLM [Code] [HF] | 🎓 Berkeley | 48.07 | 47.04 | $0.27 | 99.72 | 99.63 | 68.76 | 0.40 | 100.00 |
| 33 | RouterDC [Code] | 🎓 SUSTech | 33.75 | 32.01 | $0.07 | 39.84 | 73.00 | 49.05 | 10.75 | 85.24 |
Under review: Paix2 (#164) is temporarily removed from the ranking pending an integrity review. See #211.
🎓 Open-source 💼 Closed-source
curl -LsSf https://astral.sh/uv/install.sh | sh
cd RouterArena
uv sync
Download the dataset from HF dataset.
uv run python ./scripts/process_datasets/prep_datasets.py
In the project root, copy .env.example as .env and update the API keys in .env. This step is required only if you use our pipeline for LLM inferences.
# Example .env file
OPENAI_API_KEY=<Your-Key>
ANTHROPIC_API_KEY=<Your-Key>
# ...
See the ModelInference class for the complete list of supported providers and required environment variables. You can extend that class to support more models, or submit a GitHub issue to request support for new providers.
Follow the steps below to obtain your router's model choices for each query. Start with the sub_10 split (a 10% subset) for local testing. Once your setup works, you can evaluate:
full dataset for full local evaluation and official leaderboard submission.robustness dataset for robustness evaluation.Create a config file in ./router_inference/config/<router_name>.json. An example config file is included here.
{
"pipeline_params": {
"router_name": "your-router",
"router_cls_name": "your_router_class_name",
"models": [
"gpt-4o-mini",
"claude-3-haiku-20240307",
"gemini-2.0-flash-001"
]
}
}
For each model in your config, add an entry with the pricing per million tokens in this format at model_cost/model_cost.json:
{
"gpt-4o-mini": {
"input_token_price_per_million": 0.15,
"output_token_price_per_million": 0.6
},
}
[!NOTE] Ensure all models in your above config files are listed in
./universal_model_names.py. If you add a new model, you must also add the API inference endpoint inllm_inference/model_inference.py.
Create your own router class by inheriting from BaseRouter and implementing the _get_prediction() method. See router_inference/router/example_router.py for a complete example.
Then, modify router_inference/router/__init__.py to include your router class:
# Import your router class
from router_inference.router.my_router import MyRouter
__all__ = ["BaseRouter", "ExampleRouter", "MyRouter"]
Finally, generate the prediction file:
uv run python ./router_inference/generate_prediction_file.py your-router [sub_10|full|robustness]
[!NOTE]
- The
<your-router>argument must match your config filename (without the.jsonextension). For example, if your config file isrouter_inference/config/my-router.json, usemy-routeras the argument.- Your
_get_prediction()method must return a model name that exists in your config file'smodelslist. The base class will automatically validate this.
uv run python ./router_inference/check_config_prediction_files.py your-router [sub_10|full|robustness]
This script checks: (1) all model names are valid, (2) prediction file has correct size (809 for sub_10, 8400 for full, 420 for robustness), and (3) all entries have valid global_index, prompt, and prediction fields.
Run the inference script to make API calls for each query using the selected models:
uv run python ./llm_inference/run.py your-router
The script loads your prediction file, makes API calls using the models specified in the prediction field, and saves results incrementally. It uses cached results when available and saves progress after each query, so you can safely interrupt and resume. Results are saved to ./cached_results/ for reuse across routers.
[!NOTE]
- For robustness evaluation, we only measure the model-selection flip ratio after adding noise to the original prompt, so no additional LLM inference is required for this stage.
As the last step, run the evaluation script:
uv run python ./llm_evaluation/run.py your-router [sub_10|full|robustness]
[!TIP]
- Use
sub_10orfullto evaluate on those datasets.- Use
robustnessto run robustness-only evaluation (expects<router_name>-robustness.json).
To get your router on the leaderboard, you can open a Pull Request with your router's prediction file to trigger our automated evaluation workflow. Details are as follows:
router_inference/config/<router_name>.json - Your router configurationrouter_inference/predictions/<router_name>.json - Your prediction file with generated_result fields populatedrouter_inference/predictions/<router_name>-robustness.json - Your prediction file for robustness evaluation, no generated_result fields neededmain branch and call /evaluate in the PR comment
/evaluate in the PR comment to trigger the evaluation workflow. See an example here.The Figure below shows the evaluation pipeline.
We welcome and appreciate contributions and collaborations of any kind.
We use pre-commit to ensure a consistent coding style. You can set it up by
pip install pre-commit
pre-commit install
Before pushing your code, run the following and make sure your code passes all checks.
pre-commit run --all-files
Feel free to contact us for contributions and collaborations.
Yifan Lu (yifan.lu@rice.edu)
Rixin Liu (rixin.liu@rice.edu)
Jiarong Xing (jxing@rice.edu)
If you find our project helpful, please give us a star and cite us by:
@misc{lu2025routerarenaopenplatformcomprehensive,
title = {RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers},
author = {Yifan Lu and Rixin Liu and Jiayi Yuan and Xingqi Cui and Shenrun Zhang and Hongyi Liu and Jiarong Xing},
year = {2025},
eprint = {2510.00202},
archivePrefix= {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2510.00202}
}
Python
99.3%
RouterArena: An open framework for evaluating LLM routers with standardized datasets, metrics, an automated framework, and a live leaderboard.
See the code
RouterArena is an open evaluation platform and leaderboard for LLM routers—systems that automatically select the best model for a given query. As the LLM ecosystem diversifies with models varying in size, capability, and cost, routing has become critical for balancing performance and cost. Yet, LLM routers currently lack a standardized evaluation framework to assess how effectively they trade off accuracy, cost, and other related metrics.
RouterArena bridges this gap by providing an open evaluation platform and benchmarking framework for both open-source and commercial routers. It has the following key features:
We aim for RouterArena to serve as a foundation for the community to evaluate, understand, and advance LLM routing systems.
[!IMPORTANT] RouterArena is an evaluation-only dataset. Submissions that train, fit, or tune any router component on RouterArena data (including the label files) will be rejected, and any accepted submission found in violation will be withdrawn.
For more details, please see our website and blog.
| Rank | Router | Affiliation | Acc-Cost Arena | Accuracy | Cost/1K Queries | Optimal Selection | Optimal Cost | Optimal Accuracy | Latency | Robustness |
|---|---|---|---|---|---|---|---|---|---|---|
| 🥇 | KT-ModelRouter | 👤 @DusanBaek | 76.28 | 78.14 | $0.27 | — | — | — | — | 80.48 |
| 🥈 | Sqwish Router | 👤 @namitha-sqwish | 76.21 | 79.76 | $0.70 | 9.04 | 23.49 | 94.07 | — | 51.67 |
| 🥉 | Divyam | 👤 @samikd | 75.85 | 78.59 | $0.48 | 4.93 | 16.03 | 93.43 | — | 98.33 |
| 4 | Cross-Router | 👤 @JiaHg | 75.75 | 78.14 | $0.40 | 17.66 | 45.49 | 90.31 | — | 67.14 |
| 5 | LLM Router [PyPI] | 👤 @ypollak2 | 75.69 | 78.44 | $0.49 | — | — | — | — | 100.00 |
| 6 | R**2: Reasonometry Router | 👤 @sashakolpakov | 75.53 | 78.11 | $0.45 | — | — | — | — | 93.57 |
| 7 | SapientAI Auto Router | 💼 Publicis Sapient | 75.41 | 77.89 | $0.43 | 5.93 | 19.99 | 92.00 | — | 73.33 |
| 8 | vLLM‑SR [Code] [HF] | 🎓 vLLM SR Team | 74.86 | 77.18 | $0.42 | 16.81 | 25.10 | 89.37 | — | 67.62 |
| 9 | nadir-caliper | 👤 @doramirdor | 74.55 | 75.84 | $0.22 | — | — | — | — | 79.76 |
| 10 | AgentForge Router | 👤 @YangY-Z | 74.13 | 74.72 | $0.13 | 17.84 | 52.47 | 98.68 | — | 40.48 |
| 11 | BARouter | 👤 @zulk2002 | 73.79 | 75.72 | $0.36 | 64.41 | 67.13 | 93.84 | — | 68.81 |
| 12 | Weave Router | 🎓 Weave | 72.82 | 76.32 | $0.94 | — | — | — | — | 100.00 |
| 13 | Nadir Router | 🎓 NadirRouter | 72.29 | 75.01 | $0.68 | — | — | — | — | 25.48 |
| 14 | OrcaRouter‑Adaptive [Code] [Paper] [X] | 🎓 Continuum AI | 72.08 | 75.54 | $1.00 | — | — | — | — | 22.62 |
| 15 | Hybrid Router | 👤 @mikemao27 | 72.08 | 71.38 | $0.04 | 89.87 | 94.19 | 92.81 | — | 96.67 |
| 16 | R2-Router | 🎓 UCF | 71.60 | 71.23 | $0.06 | 24.51 | 48.70 | 99.85 | — | 45.71 |
| 17 | cruq-router | 👤 @nabaruns | 70.77 | 71.35 | $0.18 | — | — | — | — | 81.67 |
| 18 | chuzom-solo-v32 | 👤 @ypollak2 | 70.61 | 70.59 | $0.10 | — | — | — | — | 100.00 |
| 19 | Azure-Model-Router [Web] | 💼 Microsoft | 70.42 | 72.94 | $0.73 | — | — | — | — | 71.43 |
| 20 | Auto Router | 👤 @cxf2015 | 70.05 | 70.17 | $0.12 | 37.58 | 40.02 | 86.04 | — | 49.52 |
| 21 | Lynkr | 👤 @vishalveerareddy123 | 67.65 | 68.41 | $0.29 | 10.97 | 16.08 | 84.48 | — | 92.38 |
| 22 | MIRT‑BERT [Code] | 🎓 USTC | 66.89 | 66.88 | $0.15 | 3.44 | 19.62 | 78.18 | 27.03 | 61.19 |
| 23 | NIRT‑BERT [Code] | 🎓 USTC | 66.12 | 66.34 | $0.21 | 3.83 | 14.04 | 77.88 | 10.42 | 49.29 |
| 24 | AsiaInfo-Router | 👤 @Uncle-LL | 65.87 | 75.20 | $8.54 | — | — | — | — | 69.52 |
| 25 | GPT‑5 | 💼 OpenAI | 64.32 | 73.96 | $10.02 | — | — | — | — | — |
| 26 | CARROT [Code] [HF] | 🎓 UMich | 63.87 | 67.21 | $2.06 | 2.68 | 6.77 | 78.63 | 1.50 | 89.05 |
| 27 | Chayan [HF] | 🎓 Adaptive Classifier | 63.83 | 64.89 | $0.56 | 43.03 | 43.75 | 88.74 | — | — |
| 28 | RouterBench‑MLP [Code] [HF] | 🎓 Martian | 57.56 | 61.62 | $4.83 | 13.39 | 24.45 | 83.32 | 90.91 | 80.00 |
| 29 | NotDiamond | 💼 NotDiamond | 57.29 | 60.83 | $4.10 | 1.55 | 2.14 | 76.81 | — | 55.91 |
| 30 | GraphRouter [Code] | 🎓 UIUC | 57.22 | 57.00 | $0.34 | 4.73 | 38.33 | 74.25 | 2.70 | 94.29 |
| 31 | RouterBench‑KNN [Code] [HF] | 🎓 Martian | 55.48 | 58.69 | $4.27 | 13.09 | 25.49 | 78.77 | 1.33 | 83.33 |
| 32 | RouteLLM [Code] [HF] | 🎓 Berkeley | 48.07 | 47.04 | $0.27 | 99.72 | 99.63 | 68.76 | 0.40 | 100.00 |
| 33 | RouterDC [Code] | 🎓 SUSTech | 33.75 | 32.01 | $0.07 | 39.84 | 73.00 | 49.05 | 10.75 | 85.24 |
Under review: Paix2 (#164) is temporarily removed from the ranking pending an integrity review. See #211.
🎓 Open-source 💼 Closed-source
curl -LsSf https://astral.sh/uv/install.sh | sh
cd RouterArena
uv sync
Download the dataset from HF dataset.
uv run python ./scripts/process_datasets/prep_datasets.py
In the project root, copy .env.example as .env and update the API keys in .env. This step is required only if you use our pipeline for LLM inferences.
# Example .env file
OPENAI_API_KEY=<Your-Key>
ANTHROPIC_API_KEY=<Your-Key>
# ...
See the ModelInference class for the complete list of supported providers and required environment variables. You can extend that class to support more models, or submit a GitHub issue to request support for new providers.
Follow the steps below to obtain your router's model choices for each query. Start with the sub_10 split (a 10% subset) for local testing. Once your setup works, you can evaluate:
full dataset for full local evaluation and official leaderboard submission.robustness dataset for robustness evaluation.Create a config file in ./router_inference/config/<router_name>.json. An example config file is included here.
{
"pipeline_params": {
"router_name": "your-router",
"router_cls_name": "your_router_class_name",
"models": [
"gpt-4o-mini",
"claude-3-haiku-20240307",
"gemini-2.0-flash-001"
]
}
}
For each model in your config, add an entry with the pricing per million tokens in this format at model_cost/model_cost.json:
{
"gpt-4o-mini": {
"input_token_price_per_million": 0.15,
"output_token_price_per_million": 0.6
},
}
[!NOTE] Ensure all models in your above config files are listed in
./universal_model_names.py. If you add a new model, you must also add the API inference endpoint inllm_inference/model_inference.py.
Create your own router class by inheriting from BaseRouter and implementing the _get_prediction() method. See router_inference/router/example_router.py for a complete example.
Then, modify router_inference/router/__init__.py to include your router class:
# Import your router class
from router_inference.router.my_router import MyRouter
__all__ = ["BaseRouter", "ExampleRouter", "MyRouter"]
Finally, generate the prediction file:
uv run python ./router_inference/generate_prediction_file.py your-router [sub_10|full|robustness]
[!NOTE]
- The
<your-router>argument must match your config filename (without the.jsonextension). For example, if your config file isrouter_inference/config/my-router.json, usemy-routeras the argument.- Your
_get_prediction()method must return a model name that exists in your config file'smodelslist. The base class will automatically validate this.
uv run python ./router_inference/check_config_prediction_files.py your-router [sub_10|full|robustness]
This script checks: (1) all model names are valid, (2) prediction file has correct size (809 for sub_10, 8400 for full, 420 for robustness), and (3) all entries have valid global_index, prompt, and prediction fields.
Run the inference script to make API calls for each query using the selected models:
uv run python ./llm_inference/run.py your-router
The script loads your prediction file, makes API calls using the models specified in the prediction field, and saves results incrementally. It uses cached results when available and saves progress after each query, so you can safely interrupt and resume. Results are saved to ./cached_results/ for reuse across routers.
[!NOTE]
- For robustness evaluation, we only measure the model-selection flip ratio after adding noise to the original prompt, so no additional LLM inference is required for this stage.
As the last step, run the evaluation script:
uv run python ./llm_evaluation/run.py your-router [sub_10|full|robustness]
[!TIP]
- Use
sub_10orfullto evaluate on those datasets.- Use
robustnessto run robustness-only evaluation (expects<router_name>-robustness.json).
To get your router on the leaderboard, you can open a Pull Request with your router's prediction file to trigger our automated evaluation workflow. Details are as follows:
router_inference/config/<router_name>.json - Your router configurationrouter_inference/predictions/<router_name>.json - Your prediction file with generated_result fields populatedrouter_inference/predictions/<router_name>-robustness.json - Your prediction file for robustness evaluation, no generated_result fields neededmain branch and call /evaluate in the PR comment
/evaluate in the PR comment to trigger the evaluation workflow. See an example here.The Figure below shows the evaluation pipeline.
We welcome and appreciate contributions and collaborations of any kind.
We use pre-commit to ensure a consistent coding style. You can set it up by
pip install pre-commit
pre-commit install
Before pushing your code, run the following and make sure your code passes all checks.
pre-commit run --all-files
Feel free to contact us for contributions and collaborations.
Yifan Lu (yifan.lu@rice.edu)
Rixin Liu (rixin.liu@rice.edu)
Jiarong Xing (jxing@rice.edu)
If you find our project helpful, please give us a star and cite us by:
@misc{lu2025routerarenaopenplatformcomprehensive,
title = {RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers},
author = {Yifan Lu and Rixin Liu and Jiayi Yuan and Xingqi Cui and Shenrun Zhang and Hongyi Liu and Jiarong Xing},
year = {2025},
eprint = {2510.00202},
archivePrefix= {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2510.00202}
}
Python
99.3%