This release is a pre-release and may not be stable for production use.
A Living Benchmark for Machine Learning on Tabular Data 💫
| 🚀 Leaderboard | 📂 Example Scripts | 📊 Dataset Curation | 📄 Papers: TabArena-v0.1 · BeyondArena |
|---|
TabArena is a living benchmarking system that makes benchmarking tabular machine learning models a reliable experience. TabArena implements best practices to ensure methods are represented at their peak potential, including cross-validated ensembles, strong hyperparameter search spaces contributed by the method authors, early stopping, model refitting, parallel bagging, memory usage estimation, and more. Explore the latest results on the live leaderboard.
This single codebase powers two complementary benchmarks that share the same fitting, runner, and evaluation code:
- 🏟️ TabArena-v0.1 — the living benchmark on curated, IID tabular datasets.
- 🌍 BeyondArena — a holistic, beyond-IID benchmark spanning IID, temporal, and grouped tasks across a wide range of dataset sizes and feature dimensionalities. BeyondArena will superseed TabArena-v0.1 in the future.
Tip New here? Start with TabArena, then graduate to BeyondArena. Get your model working and competitive on TabArena's curated IID datasets first; once it holds up there, run the same code on BeyondArena to stress-test how well it generalizes beyond IID.
TabArena covers 51 curated datasets (9–30 splits each) and 27+ methods, including 10+ tabular foundation models — over 50M trained models, with all validation and test predictions cached for tuning and post-hoc ensembling. BeyondArena extends this to 142 datasets across IID, temporal, and grouped task types, spanning tiny to 1M-row datasets and low- to high-dimensional features.
⚡ Quickstart
Tip The fastest way to try TabArena end-to-end:
pip install uv
git clone https://github.com/autogluon/tabarena.git && cd tabarena
uv venv --seed --python 3.12 && source .venv/bin/activate
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark]"
python examples/benchmarking/run_quickstart_tabarena_model.py # benchmark a model that TabArena tunes
python examples/benchmarking/run_quickstart_tabarena_system.py # benchmark a system that tunes itself
TabArena ranks models (one method, tuned by TabArena under a shared protocol) and systems
(a pipeline that does its own tuning and ensembling, like AutoGluon); see
Contributing a Model or System for the difference and how to submit yours.
For other install paths (eval-only, editable AutoGluon, dependency), see Installation below.
To try BeyondArena instead, run python examples/beyondarena/run_quickstart_beyondarena_model.py
(or run_quickstart_beyondarena_system.py) with the same install.
🕹️ Use Cases
We share more details on various use cases of TabArena in our examples:
- 🌍 Benchmarking Beyond IID (BeyondArena): please refer to examples/beyondarena.
- 📊 Benchmarking Predictive Machine Learning Models and Systems: please refer to examples/benchmarking; to get yours onto the leaderboard, see Contributing a Model or System.
- 🧪 Advanced and Specialized Usage (incl. using a TabArena model directly on your own data): please refer to examples/advanced.
- 🗃️ Analysing Metadata and Meta-Learning: please refer to examples/meta.
- 📈 Generating Plots and Leaderboards: please refer to examples/plots.
- 🔁 Reproducibility: we share instructions for reproducibility in examples.
Datasets
Please refer to Data Foundry (documentation) to learn more about the datasets or to contribute data.
Contributing a Model or System
TabArena accepts two kinds of entrant: a model (one method that TabArena tunes under its shared protocol) and a system (a pipeline that owns its own preprocessing, validation, tuning and ensembling inside the budget TabArena hands it). TabArena is not a benchmarking service: evaluate your method on TabArena-Lite first, open a pull request with the template, and a maintainer verifies and re-runs it for the final entry. The details:
🧭 Model or system? — the difference, and where each lives in the code
A model is one method that TabArena tunes under its shared protocol: shared preprocessing, a validation split provided by TabArena, a search space of up to 200 configurations, bagging under the arena's official validation protocol (TabArena: 8 folds x 1 set, asserted by the context), and the default / tuned / tuned + ensembled variants on the leaderboard. A system owns its whole pipeline (preprocessing, validation, tuning, ensembling) inside the budget TabArena hands it: AutoML frameworks such as AutoGluon, TabFM+, LLM agents, hosted APIs. If you would have to invent a search space for your method, it is a model. If that makes no sense because the method searches for itself, it is a system.
| Model | System | |
|---|---|---|
| Code | packages/tabarena/src/tabarena/models/<key>/ |
packages/tabarena/src/tabarena/systems/<key>/ |
| Quick start | run_quickstart_tabarena_model.py, run_quickstart_beyondarena_model.py |
run_quickstart_tabarena_system.py, run_quickstart_beyondarena_system.py |
| Step-by-step guide | add-model skill |
add-system skill |
| Test | pytest -m models -k <Key> |
pytest tests/tabarena/systems/ |
📬 Submission process — evaluate, open a PR, verification, leaderboard update
We accept methods their authors have already evaluated with the official pipeline and confirm the results by re-running them.
- Say how your method differs from the entrants already on the leaderboard. A new version of an existing method goes into the existing folder and supersedes the old entry rather than becoming a new one.
- Integrate it following the guide above and run the quick start.
- Evaluate it yourself on TabArena-Lite (
subset="lite", the first split of every dataset) with HPO where applicable (the default plus about 25 random configurations), or on the BeyondArenacoresubset. The quick starts run the official validation protocol by default; every result records the protocol it ran under, and a run made withofficial_validation_protocol=Falseis declared in the pull request. - Open a pull request; the template asks for the expected files, the results, the hardware and the
entry-point script. You can also share the run's output directory (the
expnamefolder with theresults.pklfiles) so we can verify and integrate the results directly. - A maintainer reviews the pull request, then runs the method on the full task set on the benchmark hardware for the final entry. We are happy to help with the integration and the run.
- Maintainers verify that run against your Lite results, and the person you name in the template signs off on it, which marks the entry as verified.
- The results are processed, hosted and registered, and the pull request is merged.
- The leaderboard is regenerated from the hosted results after the merge, usually within days.
Questions go through the issue forms: one for model and system submissions, one for leaderboard or dataset questions that touch this code base. Pure leaderboard questions belong in the leaderboard's Community tab, dataset questions in Data Foundry. Anything else: mail@tabarena.ai.
More Documentation
There is no separate documentation site yet; the detailed reference lives in the repo and is written
for humans and coding agents alike. AGENTS.md covers the architecture, the core data
flow, models vs systems, entrant pools, caching, and the maintainer flows (processing and uploading
results, releasing to PyPI). The skills in .claude/skills/ are step-by-step guides:
add-model and add-system
for integrating a new entrant, benchmark-model for running
it on the benchmark cluster, upload-method and
update-leaderboard for publishing results, and
adapt-tabarena for building your own domain benchmark on
top of TabArena. The examples are the runnable tour.
🪄 Installation
Important Requires Python 3.11–3.13 and uv.
TabArena is a uv workspace; its installable
packages live under packages/ (tabarena, bencheval, tabflow_slurm). Install the tabarena
package directly from packages/tabarena with the extras you need. The --prerelease=allow flag is
required so uv resolves the pre-release dependency.
First clone the repo and create a virtual environment (one time):
git clone https://github.com/autogluon/tabarena.git
cd tabarena
uv venv --seed --python 3.12
source .venv/bin/activate
Then pick the install path that matches what you want to do:
📊 Evaluation only — leaderboards, metrics & plots, no model fitting
Loads cached results and computes/plots leaderboards & metrics (ELO, win-rates, ranks). Depends on autogluon.tabular (not the full AutoGluon meta-package) — no model-fitting libraries and no torch.
uv pip install --prerelease=allow -e "./packages/tabarena[plot]"
🚀 Benchmark — core set of models for benchmarking
Installs the core models used for standard benchmarking: tabpfn, tabicl, ebm, search_spaces, realmlp, tabdpt, tabm.
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark]"
➕ Benchmark + Extended — core models plus the extended model set
The
extendedextra is experimental and may fail to resolve or install due to incompatible version requirements across model dependencies. Use it only if you specifically need every model in a single environment; otherwise preferbenchmarkorbenchmarkplus one specific model.
Layers the extended model set (modernnca, xrfm, sap-rpt-oss, ...) on top of the core benchmark set.
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark,extended]"
To install only one extended model on top of benchmark (recommended over extended when you only need a single extra model), pass its extra by name — for example, just xrfm:
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark,xrfm]"
🛠️ Developer — editable AutoGluon + editable TabArena
Create a virtual environment in your workspace directory (it spans both repos cloned below, so .venv lives at the workspace root rather than inside either repo):
uv venv --seed --python 3.12 .venv
source .venv/bin/activate
Install editable AutoGluon and TabArena:
git clone https://github.com/autogluon/autogluon.git
./autogluon/full_install.sh
git clone https://github.com/autogluon/tabarena.git
uv pip install --prerelease=allow -e "./tabarena/packages/tabarena[benchmark]"
In PyCharm, mark
packages/tabarena/src/and eachautogluon/src/subdirectory as Sources Root so imports resolve.
🧪 PyPI — experimental pre-releases, no clone needed
tabarena and bencheval are published to PyPI as pre-releases for projects that cannot depend on git URLs, so pass --pre (pip) or --prerelease=allow (uv). The core package and [plot] are complete. Model extras whose upstream package is only available from git (tabfm, sap-rpt-oss, exaone_tabular) are empty on PyPI; the model's install hint tells you what to install by hand. The git checkout above stays the recommended install.
uv pip install --prerelease=allow "tabarena[plot]" # or: pip install --pre "tabarena[plot]"
uv pip install --prerelease=allow bencheval # leaderboard engine only
📦 Use TabArena as a dependency
Add one of the following to your project's dependencies:
# TabArena depends on a pre-release of AutoGluon, so allow pre-releases when installing
# (e.g. `uv pip install --prerelease=allow ...` or `pip install --pre ...`).
# Alternatively, pin AutoGluon to a specific pre-release (an exact `==` pin resolves a
# pre-release without the flag), e.g. add `"autogluon.tabular==1.5.1b20260626"`.
# From PyPI (experimental pre-releases; each tabarena release pins its matching bencheval):
"tabarena>=0.1.0a1"
# From git (tip of main; publishable to PyPI only as a source-only extra, see issue #495):
"tabarena @ git+https://github.com/autogluon/tabarena.git#subdirectory=packages/tabarena"
📦 TabArena Artifacts
TabArena caches predictions, results, and leaderboards as downloadable artifacts so you can reproduce or extend any analysis without re-running the benchmark.
Artifact tiers, sizes, and examples
Artifacts download to
~/.cache/tabarena/by default. Override the location with theTABARENA_CACHEenvironment variable.Raw data is ~100 GB per method type. Point
TABARENA_CACHEat a large disk before downloading it.
| Tier | Contents | Size / method | Example |
|---|---|---|---|
| Raw data | Per-child test predictions, full metadata, system info | ~100 GB | inspect_raw_data_and_verify_splits.py |
| Processed data | Minimal data for HPO simulation, portfolios, leaderboards | ~10 GB | inspect_processed_data.py |
| Results | Per-config / HPO DataFrames (test error, val error, train time, inference time) | <1 MB | run_generate_main_leaderboard.py |
| Leaderboards | Aggregated ELO, win-rate, average rank, improvability | <1 MB | — |
| Figures & Plots | Generated from results and leaderboards | — | — |
📄 Citation
If you use this code in a scientific publication, please cite the relevant paper(s): TabArena for the living IID benchmark, and BeyondArena for the beyond-IID benchmark.
TabArena
TabArena: A Living Benchmark for Machine Learning on Tabular Data Nick Erickson, Lennart Purucker, Andrej Tschalzev, David Holzmüller, Prateek Mutalik Desai, David Salinas, Frank Hutter NeurIPS 2025, Datasets and Benchmarks Track
📄 arXiv · 🎤 NeurIPS poster & video
BibTeX
The entry uses
year=2026because NeurIPS'25 proceedings are published in 2026.
@article{erickson2026tabarena,
title = {TabArena: A Living Benchmark for Machine Learning on Tabular Data},
author = {Erickson, Nick and Purucker, Lennart and Tschalzev, Andrej and Holzm{\"u}ller, David and Desai, Prateek and Salinas, David and Hutter, Frank},
journal = {Advances in Neural Information Processing Systems},
volume = {38},
year = {2026}
}
BeyondArena
Beyond IID: How General Are Tabular Foundation Models, Really? Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzmüller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Gaël Varoquaux, Frank Hutter
📄 arXiv
BibTeX
@misc{purucker2026beyondiid,
title = {Beyond IID: How General Are Tabular Foundation Models, Really?},
author = {Purucker, Lennart and Tschalzev, Andrej and Erickson, Nick and Blayer, Gioia and Holzm{\"u}ller, David and Arazi, Alan and Pfefferle, Alexander and Tajjar, Mustafa and Varoquaux, Ga{\"e}l and Hutter, Frank},
year = {2026},
eprint = {2606.30410},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2606.30410}
}
Relation to TabRepo
TabArena was built upon and now replaces TabRepo. To see details about TabRepo, the portfolio simulation repository, refer to tabrepo.md.
Research code
This repository contains research code intended for academic research and experimentation. It is not production-ready and should be reviewed, tested, and secured before use in production.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tabarena-0.1.1.dev20260918165950.tar.gz.
File metadata
- Download URL: tabarena-0.1.1.dev20260918165950.tar.gz
- Upload date:
- Size: 1.1 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.16 {"installer":{"name":"uv","version":"0.12.16","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f8a01e7d88bc989ff358504853ed5651effb1e8d1618ef9470dfd113c2b5f912
|
|
| MD5 |
0624e40209697897a286b048bd9af87e
|
|
| BLAKE2b-256 |
e7e76e2dca2c20c902db7008091af1ee743159201398525f605e98ff8e105455
|
File details
Details for the file tabarena-0.1.1.dev20260918165950-py3-none-any.whl.
File metadata
- Download URL: tabarena-0.1.1.dev20260918165950-py3-none-any.whl
- Upload date:
- Size: 1.3 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.16 {"installer":{"name":"uv","version":"0.12.16","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1848588c0f2f368936db773662a90966ee1bb429c54935959fd4d04ded5f7189
|
|
| MD5 |
db4a6525dfb80c6771bfae8421d7bce0
|
|
| BLAKE2b-256 |
ec7a34034e55079caf7edbe1cff465ea76e52aa5ce0ba27565f20a691c77421c
|