A Living Benchmark for Machine Learning on Tabular Data
307
stars
1,445
commits
Python
primary language
Sep 15, 2026
updated
| π Leaderboard | π Example Scripts | π Dataset Curation | π Papers: TabArena-v0.1 Β· BeyondArena |
|---|
TabArena is a living benchmarking system that makes benchmarking tabular machine learning models a reliable experience. TabArena implements best practices to ensure methods are represented at their peak potential, including cross-validated ensembles, strong hyperparameter search spaces contributed by the method authors, early stopping, model refitting, parallel bagging, memory usage estimation, and more. Explore the latest results on the live leaderboard.
This single codebase powers two complementary benchmarks that share the same fitting, runner, and evaluation code:
[!TIP] New here? Start with TabArena, then graduate to BeyondArena. Get your model working and competitive on TabArena's curated IID datasets first; once it holds up there, run the same code on BeyondArena to stress-test how well it generalizes beyond IID.
TabArena covers 51 curated datasets (9β30 splits each) and 27+ methods, including 10+ tabular foundation models β over 50M trained models, with all validation and test predictions cached for tuning and post-hoc ensembling. BeyondArena extends this to 142 datasets across IID, temporal, and grouped task types, spanning tiny to 1M-row datasets and low- to high-dimensional features.
[!TIP] The fastest way to try TabArena end-to-end:
pip install uv
git clone https://github.com/autogluon/tabarena.git && cd tabarena
uv venv --seed --python 3.12 && source .venv/bin/activate
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark]"
python examples/benchmarking/run_quickstart_tabarena_model.py # benchmark a model that TabArena tunes
python examples/benchmarking/run_quickstart_tabarena_system.py # benchmark a system that tunes itself
TabArena ranks models (one method, tuned by TabArena under a shared protocol) and systems
(a pipeline that does its own tuning and ensembling, like AutoGluon); see
Contributing a Model or System for the difference and how to submit yours.
For other install paths (eval-only, editable AutoGluon, dependency), see Installation below.
To try BeyondArena instead, run python examples/beyondarena/run_quickstart_beyondarena_model.py
(or run_quickstart_beyondarena_system.py) with the same install.
We share more details on various use cases of TabArena in our examples:
Please refer to Data Foundry (documentation) to learn more about the datasets or to contribute data.
TabArena accepts two kinds of entrant: a model (one method that TabArena tunes under its shared protocol) and a system (a pipeline that owns its own preprocessing, validation, tuning and ensembling inside the budget TabArena hands it). TabArena is not a benchmarking service: evaluate your method on TabArena-Lite first, open a pull request with the template, and a maintainer verifies and re-runs it for the final entry. The details:
A model is one method that TabArena tunes under its shared protocol: shared preprocessing, a validation split provided by TabArena, a search space of up to 200 configurations, bagging, and the default / tuned / tuned + ensembled variants on the leaderboard. A system owns its whole pipeline (preprocessing, validation, tuning, ensembling) inside the budget TabArena hands it: AutoML frameworks such as AutoGluon, TabFM+, LLM agents, hosted APIs. If you would have to invent a search space for your method, it is a model. If that makes no sense because the method searches for itself, it is a system.
| Model | System | |
|---|---|---|
| Code | packages/tabarena/src/tabarena/models/<key>/ | packages/tabarena/src/tabarena/systems/<key>/ |
| Quick start | run_quickstart_tabarena_model.py, run_quickstart_beyondarena_model.py | run_quickstart_tabarena_system.py, run_quickstart_beyondarena_system.py |
| Step-by-step guide | add-model skill | add-system skill |
| Test | pytest -m models -k <Key> | pytest tests/tabarena/systems/ |
We accept methods their authors have already evaluated with the official pipeline and confirm the results by re-running them.
subset="lite", the first split of every dataset) with HPO
where applicable (the default plus about 25 random configurations), or on the BeyondArena core
subset.expname folder with the
results.pkl files) so we can verify and integrate the results directly.Questions go through the issue forms: one for model and system submissions, one for leaderboard or dataset questions that touch this code base. Pure leaderboard questions belong in the leaderboard's Community tab, dataset questions in Data Foundry. Anything else: mail@tabarena.ai.
There is no separate documentation site yet; the detailed reference lives in the repo and is written
for humans and coding agents alike. AGENTS.md covers the architecture, the core data
flow, models vs systems, entrant pools, caching, and the maintainer flows (processing and uploading
results, releasing to PyPI). The skills in .claude/skills/ are step-by-step guides:
add-model and add-system
for integrating a new entrant, benchmark-model for running
it on the benchmark cluster, upload-method and
update-leaderboard for publishing results, and
adapt-tabarena for building your own domain benchmark on
top of TabArena. The examples are the runnable tour.
[!IMPORTANT] Requires Python 3.11β3.13 and uv.
TabArena is a uv workspace; its installable
packages live under packages/ (tabarena, bencheval, tabflow_slurm). Install the tabarena
package directly from packages/tabarena with the extras you need. The --prerelease=allow flag is
required so uv resolves the pre-release dependency.
First clone the repo and create a virtual environment (one time):
git clone https://github.com/autogluon/tabarena.git
cd tabarena
uv venv --seed --python 3.12
source .venv/bin/activate
Then pick the install path that matches what you want to do:
Loads cached results and computes/plots leaderboards & metrics (ELO, win-rates, ranks). Depends on autogluon.tabular (not the full AutoGluon meta-package) β no model-fitting libraries and no torch.
uv pip install --prerelease=allow -e "./packages/tabarena[plot]"
Installs the core models used for standard benchmarking: tabpfn, tabicl, ebm, search_spaces, realmlp, tabdpt, tabm.
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark]"
The
extendedextra is experimental and may fail to resolve or install due to incompatible version requirements across model dependencies. Use it only if you specifically need every model in a single environment; otherwise preferbenchmarkorbenchmarkplus one specific model.
Layers the extended model set (modernnca, xrfm, sap-rpt-oss, ...) on top of the core benchmark set.
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark,extended]"
To install only one extended model on top of benchmark (recommended over extended when you only need a single extra model), pass its extra by name β for example, just xrfm:
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark,xrfm]"
Create a virtual environment in your workspace directory (it spans both repos cloned below, so .venv lives at the workspace root rather than inside either repo):
uv venv --seed --python 3.12 .venv
source .venv/bin/activate
Install editable AutoGluon and TabArena:
git clone https://github.com/autogluon/autogluon.git
./autogluon/full_install.sh
git clone https://github.com/autogluon/tabarena.git
uv pip install --prerelease=allow -e "./tabarena/packages/tabarena[benchmark]"
In PyCharm, mark
packages/tabarena/src/and eachautogluon/src/subdirectory as Sources Root so imports resolve.
tabarena and bencheval are published to PyPI as pre-releases for projects that cannot depend on git URLs, so pass --pre (pip) or --prerelease=allow (uv). The core package and [plot] are complete. Model extras whose upstream package is only available from git (tabfm, sap-rpt-oss, exaone_tabular) are empty on PyPI; the model's install hint tells you what to install by hand. The git checkout above stays the recommended install.
uv pip install --prerelease=allow "tabarena[plot]" # or: pip install --pre "tabarena[plot]"
uv pip install --prerelease=allow bencheval # leaderboard engine only
Add one of the following to your project's dependencies:
# TabArena depends on a pre-release of AutoGluon, so allow pre-releases when installing
# (e.g. `uv pip install --prerelease=allow ...` or `pip install --pre ...`).
# Alternatively, pin AutoGluon to a specific pre-release (an exact `==` pin resolves a
# pre-release without the flag), e.g. add `"autogluon.tabular==1.5.1b20260626"`.
# From PyPI (experimental pre-releases; each tabarena release pins its matching bencheval):
"tabarena>=0.1.0a1"
# From git (tip of main; publishable to PyPI only as a source-only extra, see issue #495):
"tabarena @ git+https://github.com/autogluon/tabarena.git#subdirectory=packages/tabarena"
TabArena caches predictions, results, and leaderboards as downloadable artifacts so you can reproduce or extend any analysis without re-running the benchmark.
Artifacts download to
~/.cache/tabarena/by default. Override the location with theTABARENA_CACHEenvironment variable.Raw data is ~100 GB per method type. Point
TABARENA_CACHEat a large disk before downloading it.
| Tier | Contents | Size / method | Example |
|---|---|---|---|
| Raw data | Per-child test predictions, full metadata, system info | ~100 GB | inspect_raw_data_and_verify_splits.py |
| Processed data | Minimal data for HPO simulation, portfolios, leaderboards | ~10 GB | inspect_processed_data.py |
| Results | Per-config / HPO DataFrames (test error, val error, train time, inference time) | <1 MB | run_generate_main_leaderboard.py |
| Leaderboards | Aggregated ELO, win-rate, average rank, improvability | <1 MB | β |
| Figures & Plots | Generated from results and leaderboards | β | β |
If you use this code in a scientific publication, please cite the relevant paper(s): TabArena for the living IID benchmark, and BeyondArena for the beyond-IID benchmark.
TabArena: A Living Benchmark for Machine Learning on Tabular Data Nick Erickson, Lennart Purucker, Andrej Tschalzev, David HolzmΓΌller, Prateek Mutalik Desai, David Salinas, Frank Hutter NeurIPS 2025, Datasets and Benchmarks Track
π arXiv Β· π€ NeurIPS poster & video
The entry uses
year=2026because NeurIPS'25 proceedings are published in 2026.
@article{erickson2026tabarena,
title = {TabArena: A Living Benchmark for Machine Learning on Tabular Data},
author = {Erickson, Nick and Purucker, Lennart and Tschalzev, Andrej and Holzm{\"u}ller, David and Desai, Prateek and Salinas, David and Hutter, Frank},
journal = {Advances in Neural Information Processing Systems},
volume = {38},
year = {2026}
}
Beyond IID: How General Are Tabular Foundation Models, Really? Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David HolzmΓΌller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, GaΓ«l Varoquaux, Frank Hutter
π arXiv
@misc{purucker2026beyondiid,
title = {Beyond IID: How General Are Tabular Foundation Models, Really?},
author = {Purucker, Lennart and Tschalzev, Andrej and Erickson, Nick and Blayer, Gioia and Holzm{\"u}ller, David and Arazi, Alan and Pfefferle, Alexander and Tajjar, Mustafa and Varoquaux, Ga{\"e}l and Hutter, Frank},
year = {2026},
eprint = {2606.30410},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2606.30410}
}
TabArena was built upon and now replaces TabRepo. To see details about TabRepo, the portfolio simulation repository, refer to tabrepo.md.
This repository contains research code intended for academic research and experimentation. It is not production-ready and should be reviewed, tested, and secured before use in production.
Python
99.7%
A Living Benchmark for Machine Learning on Tabular Data
307
stars
1,445
commits
Python
primary language
Sep 15, 2026
updated
| π Leaderboard | π Example Scripts | π Dataset Curation | π Papers: TabArena-v0.1 Β· BeyondArena |
|---|
TabArena is a living benchmarking system that makes benchmarking tabular machine learning models a reliable experience. TabArena implements best practices to ensure methods are represented at their peak potential, including cross-validated ensembles, strong hyperparameter search spaces contributed by the method authors, early stopping, model refitting, parallel bagging, memory usage estimation, and more. Explore the latest results on the live leaderboard.
This single codebase powers two complementary benchmarks that share the same fitting, runner, and evaluation code:
[!TIP] New here? Start with TabArena, then graduate to BeyondArena. Get your model working and competitive on TabArena's curated IID datasets first; once it holds up there, run the same code on BeyondArena to stress-test how well it generalizes beyond IID.
TabArena covers 51 curated datasets (9β30 splits each) and 27+ methods, including 10+ tabular foundation models β over 50M trained models, with all validation and test predictions cached for tuning and post-hoc ensembling. BeyondArena extends this to 142 datasets across IID, temporal, and grouped task types, spanning tiny to 1M-row datasets and low- to high-dimensional features.
[!TIP] The fastest way to try TabArena end-to-end:
pip install uv
git clone https://github.com/autogluon/tabarena.git && cd tabarena
uv venv --seed --python 3.12 && source .venv/bin/activate
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark]"
python examples/benchmarking/run_quickstart_tabarena_model.py # benchmark a model that TabArena tunes
python examples/benchmarking/run_quickstart_tabarena_system.py # benchmark a system that tunes itself
TabArena ranks models (one method, tuned by TabArena under a shared protocol) and systems
(a pipeline that does its own tuning and ensembling, like AutoGluon); see
Contributing a Model or System for the difference and how to submit yours.
For other install paths (eval-only, editable AutoGluon, dependency), see Installation below.
To try BeyondArena instead, run python examples/beyondarena/run_quickstart_beyondarena_model.py
(or run_quickstart_beyondarena_system.py) with the same install.
We share more details on various use cases of TabArena in our examples:
Please refer to Data Foundry (documentation) to learn more about the datasets or to contribute data.
TabArena accepts two kinds of entrant: a model (one method that TabArena tunes under its shared protocol) and a system (a pipeline that owns its own preprocessing, validation, tuning and ensembling inside the budget TabArena hands it). TabArena is not a benchmarking service: evaluate your method on TabArena-Lite first, open a pull request with the template, and a maintainer verifies and re-runs it for the final entry. The details:
A model is one method that TabArena tunes under its shared protocol: shared preprocessing, a validation split provided by TabArena, a search space of up to 200 configurations, bagging, and the default / tuned / tuned + ensembled variants on the leaderboard. A system owns its whole pipeline (preprocessing, validation, tuning, ensembling) inside the budget TabArena hands it: AutoML frameworks such as AutoGluon, TabFM+, LLM agents, hosted APIs. If you would have to invent a search space for your method, it is a model. If that makes no sense because the method searches for itself, it is a system.
| Model | System | |
|---|---|---|
| Code | packages/tabarena/src/tabarena/models/<key>/ | packages/tabarena/src/tabarena/systems/<key>/ |
| Quick start | run_quickstart_tabarena_model.py, run_quickstart_beyondarena_model.py | run_quickstart_tabarena_system.py, run_quickstart_beyondarena_system.py |
| Step-by-step guide | add-model skill | add-system skill |
| Test | pytest -m models -k <Key> | pytest tests/tabarena/systems/ |
We accept methods their authors have already evaluated with the official pipeline and confirm the results by re-running them.
subset="lite", the first split of every dataset) with HPO
where applicable (the default plus about 25 random configurations), or on the BeyondArena core
subset.expname folder with the
results.pkl files) so we can verify and integrate the results directly.Questions go through the issue forms: one for model and system submissions, one for leaderboard or dataset questions that touch this code base. Pure leaderboard questions belong in the leaderboard's Community tab, dataset questions in Data Foundry. Anything else: mail@tabarena.ai.
There is no separate documentation site yet; the detailed reference lives in the repo and is written
for humans and coding agents alike. AGENTS.md covers the architecture, the core data
flow, models vs systems, entrant pools, caching, and the maintainer flows (processing and uploading
results, releasing to PyPI). The skills in .claude/skills/ are step-by-step guides:
add-model and add-system
for integrating a new entrant, benchmark-model for running
it on the benchmark cluster, upload-method and
update-leaderboard for publishing results, and
adapt-tabarena for building your own domain benchmark on
top of TabArena. The examples are the runnable tour.
[!IMPORTANT] Requires Python 3.11β3.13 and uv.
TabArena is a uv workspace; its installable
packages live under packages/ (tabarena, bencheval, tabflow_slurm). Install the tabarena
package directly from packages/tabarena with the extras you need. The --prerelease=allow flag is
required so uv resolves the pre-release dependency.
First clone the repo and create a virtual environment (one time):
git clone https://github.com/autogluon/tabarena.git
cd tabarena
uv venv --seed --python 3.12
source .venv/bin/activate
Then pick the install path that matches what you want to do:
Loads cached results and computes/plots leaderboards & metrics (ELO, win-rates, ranks). Depends on autogluon.tabular (not the full AutoGluon meta-package) β no model-fitting libraries and no torch.
uv pip install --prerelease=allow -e "./packages/tabarena[plot]"
Installs the core models used for standard benchmarking: tabpfn, tabicl, ebm, search_spaces, realmlp, tabdpt, tabm.
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark]"
The
extendedextra is experimental and may fail to resolve or install due to incompatible version requirements across model dependencies. Use it only if you specifically need every model in a single environment; otherwise preferbenchmarkorbenchmarkplus one specific model.
Layers the extended model set (modernnca, xrfm, sap-rpt-oss, ...) on top of the core benchmark set.
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark,extended]"
To install only one extended model on top of benchmark (recommended over extended when you only need a single extra model), pass its extra by name β for example, just xrfm:
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark,xrfm]"
Create a virtual environment in your workspace directory (it spans both repos cloned below, so .venv lives at the workspace root rather than inside either repo):
uv venv --seed --python 3.12 .venv
source .venv/bin/activate
Install editable AutoGluon and TabArena:
git clone https://github.com/autogluon/autogluon.git
./autogluon/full_install.sh
git clone https://github.com/autogluon/tabarena.git
uv pip install --prerelease=allow -e "./tabarena/packages/tabarena[benchmark]"
In PyCharm, mark
packages/tabarena/src/and eachautogluon/src/subdirectory as Sources Root so imports resolve.
tabarena and bencheval are published to PyPI as pre-releases for projects that cannot depend on git URLs, so pass --pre (pip) or --prerelease=allow (uv). The core package and [plot] are complete. Model extras whose upstream package is only available from git (tabfm, sap-rpt-oss, exaone_tabular) are empty on PyPI; the model's install hint tells you what to install by hand. The git checkout above stays the recommended install.
uv pip install --prerelease=allow "tabarena[plot]" # or: pip install --pre "tabarena[plot]"
uv pip install --prerelease=allow bencheval # leaderboard engine only
Add one of the following to your project's dependencies:
# TabArena depends on a pre-release of AutoGluon, so allow pre-releases when installing
# (e.g. `uv pip install --prerelease=allow ...` or `pip install --pre ...`).
# Alternatively, pin AutoGluon to a specific pre-release (an exact `==` pin resolves a
# pre-release without the flag), e.g. add `"autogluon.tabular==1.5.1b20260626"`.
# From PyPI (experimental pre-releases; each tabarena release pins its matching bencheval):
"tabarena>=0.1.0a1"
# From git (tip of main; publishable to PyPI only as a source-only extra, see issue #495):
"tabarena @ git+https://github.com/autogluon/tabarena.git#subdirectory=packages/tabarena"
TabArena caches predictions, results, and leaderboards as downloadable artifacts so you can reproduce or extend any analysis without re-running the benchmark.
Artifacts download to
~/.cache/tabarena/by default. Override the location with theTABARENA_CACHEenvironment variable.Raw data is ~100 GB per method type. Point
TABARENA_CACHEat a large disk before downloading it.
| Tier | Contents | Size / method | Example |
|---|---|---|---|
| Raw data | Per-child test predictions, full metadata, system info | ~100 GB | inspect_raw_data_and_verify_splits.py |
| Processed data | Minimal data for HPO simulation, portfolios, leaderboards | ~10 GB | inspect_processed_data.py |
| Results | Per-config / HPO DataFrames (test error, val error, train time, inference time) | <1 MB | run_generate_main_leaderboard.py |
| Leaderboards | Aggregated ELO, win-rate, average rank, improvability | <1 MB | β |
| Figures & Plots | Generated from results and leaderboards | β | β |
If you use this code in a scientific publication, please cite the relevant paper(s): TabArena for the living IID benchmark, and BeyondArena for the beyond-IID benchmark.
TabArena: A Living Benchmark for Machine Learning on Tabular Data Nick Erickson, Lennart Purucker, Andrej Tschalzev, David HolzmΓΌller, Prateek Mutalik Desai, David Salinas, Frank Hutter NeurIPS 2025, Datasets and Benchmarks Track
π arXiv Β· π€ NeurIPS poster & video
The entry uses
year=2026because NeurIPS'25 proceedings are published in 2026.
@article{erickson2026tabarena,
title = {TabArena: A Living Benchmark for Machine Learning on Tabular Data},
author = {Erickson, Nick and Purucker, Lennart and Tschalzev, Andrej and Holzm{\"u}ller, David and Desai, Prateek and Salinas, David and Hutter, Frank},
journal = {Advances in Neural Information Processing Systems},
volume = {38},
year = {2026}
}
Beyond IID: How General Are Tabular Foundation Models, Really? Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David HolzmΓΌller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, GaΓ«l Varoquaux, Frank Hutter
π arXiv
@misc{purucker2026beyondiid,
title = {Beyond IID: How General Are Tabular Foundation Models, Really?},
author = {Purucker, Lennart and Tschalzev, Andrej and Erickson, Nick and Blayer, Gioia and Holzm{\"u}ller, David and Arazi, Alan and Pfefferle, Alexander and Tajjar, Mustafa and Varoquaux, Ga{\"e}l and Hutter, Frank},
year = {2026},
eprint = {2606.30410},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2606.30410}
}
TabArena was built upon and now replaces TabRepo. To see details about TabRepo, the portfolio simulation repository, refer to tabrepo.md.
This repository contains research code intended for academic research and experimentation. It is not production-ready and should be reviewed, tested, and secured before use in production.
Python
99.7%