Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

    TabArena Logo

A Living Benchmark for Machine Learning on Tabular Data 💫


🚀 Leaderboard 📂 Example Scripts 📊 Dataset Curation 📄 Papers: TabArena-v0.1 · BeyondArena

TabArena is a living benchmarking system that makes benchmarking tabular machine learning models a reliable experience. TabArena implements best practices to ensure methods are represented at their peak potential, including cross-validated ensembles, strong hyperparameter search spaces contributed by the method authors, early stopping, model refitting, parallel bagging, memory usage estimation, and more. Explore the latest results on the live leaderboard.

This single codebase powers two complementary benchmarks that share the same fitting, runner, and evaluation code:

  • 🏟️ TabArena-v0.1 — the living benchmark on curated, IID tabular datasets.
  • 🌍 BeyondArena — a holistic, beyond-IID benchmark spanning IID, temporal, and grouped tasks across a wide range of dataset sizes and feature dimensionalities. BeyondArena will superseed TabArena-v0.1 in the future.

Tip New here? Start with TabArena, then graduate to BeyondArena. Get your model working and competitive on TabArena's curated IID datasets first; once it holds up there, run the same code on BeyondArena to stress-test how well it generalizes beyond IID.

TabArena covers 51 curated datasets (9–30 splits each) and 27+ methods, including 10+ tabular foundation models — over 50M trained models, with all validation and test predictions cached for tuning and post-hoc ensembling. BeyondArena extends this to 142 datasets across IID, temporal, and grouped task types, spanning tiny to 1M-row datasets and low- to high-dimensional features.

⚡ Quickstart

Tip The fastest way to try TabArena end-to-end:

pip install uv
git clone https://github.com/autogluon/tabarena.git && cd tabarena
uv venv --seed --python 3.12 && source .venv/bin/activate
uv pip install --prerelease=allow -e "./packages/tabarena[benchmark]"
python examples/benchmarking/run_quickstart_tabarena.py

For other install paths (eval-only, editable AutoGluon, dependency), see Installation below. To try BeyondArena instead, run python examples/beyondarena/run_quickstart_beyondarena.py with the same install.

🕹️ Use Cases

We share more details on various use cases of TabArena in our examples:

Datasets

Please refer to our dataset curation repository to learn more about or contributed data!

More Documentation

TabArena code is currently being polished. Detailed Documentation for TabArena will be available soon.

🪄 Installation

Important Requires Python 3.11–3.13 and uv.

TabArena is a uv workspace; its installable packages live under packages/ (tabarena, bencheval, tabflow_slurm). Install the tabarena package directly from packages/tabarena with the extras you need. The --prerelease=allow flag is required so uv resolves the pre-release dependency.

First clone the repo and create a virtual environment (one time):

git clone https://github.com/autogluon/tabarena.git
cd tabarena
uv venv --seed --python 3.12
source .venv/bin/activate

Then pick the install path that matches what you want to do:

📊 Evaluation only — leaderboards, metrics & plots, no model fitting

Loads cached results and computes/plots leaderboards & metrics (ELO, win-rates, ranks). Depends on autogluon.tabular (not the full AutoGluon meta-package) — no model-fitting libraries and no torch.

uv pip install --prerelease=allow -e "./packages/tabarena[plot]"
🚀 Benchmark — core set of models for benchmarking

Installs the core models used for standard benchmarking: tabpfn, tabicl, ebm, search_spaces, realmlp, tabdpt, tabm.

uv pip install --prerelease=allow -e "./packages/tabarena[benchmark]"
➕ Benchmark + Extended — core models plus the extended model set

The extended extra is experimental and may fail to resolve or install due to incompatible version requirements across model dependencies. Use it only if you specifically need every model in a single environment; otherwise prefer benchmark or benchmark plus one specific model.

Layers the extended model set (modernnca, xrfm, sap-rpt-oss, ...) on top of the core benchmark set.

uv pip install --prerelease=allow -e "./packages/tabarena[benchmark,extended]"

To install only one extended model on top of benchmark (recommended over extended when you only need a single extra model), pass its extra by name — for example, just xrfm:

uv pip install --prerelease=allow -e "./packages/tabarena[benchmark,xrfm]"
🛠️ Developer — editable AutoGluon + editable TabArena

Create a virtual environment in your workspace directory (it spans both repos cloned below, so .venv lives at the workspace root rather than inside either repo):

uv venv --seed --python 3.12 .venv
source .venv/bin/activate

Install editable AutoGluon and TabArena:

git clone https://github.com/autogluon/autogluon.git
./autogluon/full_install.sh

git clone https://github.com/autogluon/tabarena.git
uv pip install --prerelease=allow -e "./tabarena/packages/tabarena[benchmark]"

In PyCharm, mark packages/tabarena/src/ and each autogluon/src/ subdirectory as Sources Root so imports resolve.

🧪 PyPI — experimental pre-releases, no clone needed

tabarena and bencheval are published to PyPI as pre-releases for projects that cannot depend on git URLs, so pass --pre (pip) or --prerelease=allow (uv). The core package and [plot] are complete. Model extras whose upstream package is only available from git (tabfm, sap-rpt-oss, exaone_tabular) are empty on PyPI; the model's install hint tells you what to install by hand. The git checkout above stays the recommended install.

uv pip install --prerelease=allow "tabarena[plot]"   # or: pip install --pre "tabarena[plot]"
uv pip install --prerelease=allow bencheval          # leaderboard engine only
📦 Use TabArena as a dependency

Add one of the following to your project's dependencies:

# TabArena depends on a pre-release of AutoGluon, so allow pre-releases when installing
# (e.g. `uv pip install --prerelease=allow ...` or `pip install --pre ...`).
# Alternatively, pin AutoGluon to a specific pre-release (an exact `==` pin resolves a
# pre-release without the flag), e.g. add `"autogluon.tabular==1.5.1b20260626"`.

# From PyPI (experimental pre-releases; each tabarena release pins its matching bencheval):
"tabarena>=0.1.0a1"
# From git (tip of main; publishable to PyPI only as a source-only extra, see issue #495):
"tabarena @ git+https://github.com/autogluon/tabarena.git#subdirectory=packages/tabarena"

📦 TabArena Artifacts

TabArena caches predictions, results, and leaderboards as downloadable artifacts so you can reproduce or extend any analysis without re-running the benchmark.

Artifact tiers, sizes, and examples

Artifacts download to ~/.cache/tabarena/ by default. Override the location with the TABARENA_CACHE environment variable.

Raw data is ~100 GB per method type. Point TABARENA_CACHE at a large disk before downloading it.

Tier Contents Size / method Example
Raw data Per-child test predictions, full metadata, system info ~100 GB inspect_raw_data_and_verify_splits.py
Processed data Minimal data for HPO simulation, portfolios, leaderboards ~10 GB inspect_processed_data.py
Results Per-config / HPO DataFrames (test error, val error, train time, inference time) <1 MB run_generate_main_leaderboard.py
Leaderboards Aggregated ELO, win-rate, average rank, improvability <1 MB
Figures & Plots Generated from results and leaderboards

📄 Citation

If you use this code in a scientific publication, please cite the relevant paper(s): TabArena for the living IID benchmark, and BeyondArena for the beyond-IID benchmark.

TabArena

TabArena: A Living Benchmark for Machine Learning on Tabular Data Nick Erickson, Lennart Purucker, Andrej Tschalzev, David Holzmüller, Prateek Mutalik Desai, David Salinas, Frank Hutter NeurIPS 2025, Datasets and Benchmarks Track

📄 arXiv · 🎤 NeurIPS poster & video

BibTeX

The entry uses year=2026 because NeurIPS'25 proceedings are published in 2026.

@article{erickson2026tabarena,
  title   = {TabArena: A Living Benchmark for Machine Learning on Tabular Data},
  author  = {Erickson, Nick and Purucker, Lennart and Tschalzev, Andrej and Holzm{\"u}ller, David and Desai, Prateek and Salinas, David and Hutter, Frank},
  journal = {Advances in Neural Information Processing Systems},
  volume  = {38},
  year    = {2026}
}

BeyondArena

Beyond IID: How General Are Tabular Foundation Models, Really? Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzmüller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Gaël Varoquaux, Frank Hutter

📄 arXiv

BibTeX
@misc{purucker2026beyondiid,
  title         = {Beyond IID: How General Are Tabular Foundation Models, Really?},
  author        = {Purucker, Lennart and Tschalzev, Andrej and Erickson, Nick and Blayer, Gioia and Holzm{\"u}ller, David and Arazi, Alan and Pfefferle, Alexander and Tajjar, Mustafa and Varoquaux, Ga{\"e}l and Hutter, Frank},
  year          = {2026},
  eprint        = {2606.30410},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2606.30410}
}

Relation to TabRepo

TabArena was built upon and now replaces TabRepo. To see details about TabRepo, the portfolio simulation repository, refer to tabrepo.md.

Research code

This repository contains research code intended for academic research and experimentation. It is not production-ready and should be reviewed, tested, and secured before use in production.

Release files for tabarena 0.1.1.dev20260903151427

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tabarena 0.1.1.dev20260903151427
File Size Uploaded
tabarena-0.1.1.dev20260903151427.tar.gz 926.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tabarena 0.1.1.dev20260903151427
File Interpreter ABI Platform
tabarena-0.1.1.dev20260903151427-py3-none-any.whl Python 3 none any Details

Total release size:2.0 MB

Release files / tabarena-0.1.1.dev20260903151427.tar.gz

Download URL tabarena-0.1.1.dev20260903151427.tar.gz
Size 926.6 kB
Tags Source
SHA-256 checksum
How to use checksums
519f3015c4bc2e3b80c26c5948ddc1148bb6a18fd5448fc4d60df9f7841f8908
BLAKE2b-256 checksum
How to use checksums
dde2b498b1c3119f36ae482df69efcc1a0f921a568d4462167da83587ce23161
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.9 {"installer":{"name":"uv","version":"0.12.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / tabarena-0.1.1.dev20260903151427-py3-none-any.whl

Download URL tabarena-0.1.1.dev20260903151427-py3-none-any.whl
Size 1.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
3c84cb456f0e7ea97f624657752ff5068fcabe8ea31b451c2160f5b1f2ec19e5
BLAKE2b-256 checksum
How to use checksums
b89c646667b003a2b1476230c41e89e64d9af4f5aea4e411a792673720012d43
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.9 {"installer":{"name":"uv","version":"0.12.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.1.0

2 release files

0.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page