Skip to main content

ScAn-Bench: Surrogate Benchmarks for Scaling Analysis

PyPI version Python versions License Tests

Training runs, empirical Pareto fronts and power-law fits for the OpenCLIP, LLM and TabPFN benchmarks

The three ScAn-Bench suites. Grey: training runs. Orange: empirical compute-optimal front. Dotted: power-law fit.

Scaling studies are expensive, and every new method pays for them again. To validate a scaling analysis method, groups train hundreds of models, burn thousands of GPU-hours (and the energy that comes with them), then throw the sweep away. The next paper repeats it from scratch.

Existing benchmarks also only cover half of the problem. They either tune hyperparameters such as learning rate and batch size for one fixed architecture, or scale the architecture while fixing those hyperparameters to heuristic values. Scaling analysis needs both to vary together.

ScAn-Bench does the expensive part once. We trained 5,352 configurations across three model families, jointly varying model scale, data scale and training hyperparameters over 2–3 orders of magnitude of compute. Surrogate models fitted to these runs let you query any configuration in the search space, at any scale, in under 20 seconds, with no GPU required.

Benchmark Model family Configurations Compute range (FLOPs) Targets
VLM OpenCLIP 1,535 1014 – 1016 val/test loss, divergence report
LLM Decoder-only Transformer 1,194 1016 – 1019 val/test loss
TabPFN Encoder-only Tabular foundation model 2,623 1011 – 1014 prior validation loss

For a general overview of the paper, repository structure and artifact maps refer to PaperOverview.

To create a new benchmark based on this framework refer to BenchmarkCreation.

ScAn-Bench Repository

SDK Installation

We recommend using a conda environment for installing the SDK:

pip install scan-bench

Install pytorch with CUDA, if you want to utilize the GPU.

Quick start

from scan_bench import TabPFNBenchmark, TabPFNConfig, TabPFNTarget, PerformancePredictorType

bench = TabPFNBenchmark(
    target=TabPFNTarget.VAL_LOSS,
    predictor_type=PerformancePredictorType.ENSEMBLE_XGB,  # or TABPFN, ENSEMBLE_LIGHTGBM, ENSEMBLE_MIX, AUTOGLUON
    device="auto",                                         # or "cpu", "cuda"
)

small = TabPFNConfig(total_cells=2**20, effective_batch_size=16, lr=1e-4,
                     max_features=32, embedding_size=4, num_layers=2, num_datapoints_max=128)
large = TabPFNConfig(total_cells=2**30, effective_batch_size=64, lr=1e-3,
                     max_features=64, embedding_size=64, num_layers=8, num_datapoints_max=256)

print(bench.query(small))                  # {"predictions": {"mean", "uncertainty"}, "model_stats": {"flops", ...}}
print(bench.query_many([small, large]))    # batched, one result per config
print(bench.flops(large), bench.model_params(large))

Package usage

Refer to VLM API, LLM API and TabPFN API for API usages.

Local development and experiment reproduction

For local development and experiment reproduction, use the cloned repository. The training, evaluation, and plotting shell scripts described below are source-tree workflows and should be run from the repository root.

Surrogate training and evaluation

To train and get the performance results for the surrogate benchmarks, run the provided shell scripts. Change DEVICE to 'cuda' in train_surrogates.sh to use the GPU.

VLM pipeline

VLM performance predictor surrogate
bash scan_bench/vlm/performance_surrogate/train/train_surrogates.sh
VLM divergences predictor surrogate
bash scan_bench/vlm/divergence_surrogate/train.sh

LLM pipeline

bash scan_bench/llm/train_surrogates.sh

TabPFN pipeline

bash scan_bench/tabpfn/performance_surrogate/train/train_surrogates.sh

Results

After running the surrogate training scripts, a results directory is created under the corresponding directory for each model family. The generated subdirectories contain JSON files with the evaluation results for the trained surrogate models. These JSON outputs include the metrics reported below and are used to construct the corresponding tables in the paper.

For example:

scan_bench/vlm/performance_surrogate/results/

scan_bench/llm/results/

scan_bench/tabpfn/performance_surrogate/results/

Surrogate Performance (VLM)

Surrogate RMSE ↓ MAE ↓ MDAE ↓ MARPD ↓ R² ↑ R ↑ Corr. ↑
TabPFN 0.21 0.06 0.02 2.69 0.96 0.98 0.98
AutoGluon 0.26 0.12 0.06 5.79 0.94 0.97 0.98
XGB 0.35 0.20 0.12 9.11 0.90 0.95 0.96
Mix 0.37 0.21 0.14 10.11 0.88 0.94 0.95
LGB 0.42 0.29 0.24 14.88 0.86 0.94 0.95

Surrogate Performance (LLM)

Surrogate RMSE ↓ MAE ↓ MDAE ↓ MARPD ↓ R² ↑ R ↑ Corr. ↑
TabPFN 0.27 0.08 0.01 3.02 0.77 0.88 0.95
AutoGluon 0.30 0.11 0.02 4.66 0.72 0.85 0.91
XGB 0.35 0.15 0.04 6.80 0.63 0.80 0.86
Mix 0.33 0.17 0.07 7.65 0.66 0.82 0.86
LGB 0.33 0.17 0.08 7.97 0.66 0.83 0.87

Data

The repository includes pre-collected configuration-performance datasets used to train the surrogate models.

VLM dataset summary

Quantity Count Description
Training configurations 1,535 Total number of collected VLM training configurations.
Failed configurations 134 Configurations that diverged during training.
Successful configurations 1366 Configurations that completed training successfully.
Collected checkpoint rows 8,024 Total checkpoint rows collected across all runs.
Successful checkpoint rows 7,701 Checkpoints from successful runs

LLM dataset summary

Quantity Count Description
Training configurations 1,194 Total number of collected LLM training configurations.
Collected checkpoint rows 4,524 Total checkpoints collected across successful and failed runs.
Performance-surrogate rows 4,524 Checkpoints from successful runs used for performance prediction.

TabPFN dataset summary

Quantity Count Description
Training configurations 2,623 Total number of collected TabPFN pretraining configurations (one row per configuration).

Each configuration varies the hyperparameters (lr, effective_batch_size) and the scale parameters (total_cells, embedding_size, num_layers, max_features, num_datapoints_max); see TabPFN search space. Available target is the prior validation loss (val/val_loss).

Data locations

The table below shows where the data is located:

Dataset Path Description
VLM performance data scan_bench/vlm/performance_surrogate/splits Training and test splits for VLM performance surrogate modeling.
VLM divergence data scan_bench/vlm/divergence_surrogate/splits Configuration-level data for predicting failed (diverged) configurations.
LLM performance data scan_bench/llm/splits Configuration-performance datasets for LLM surrogate training.
TabPFN performance data scan_bench/tabpfn/performance_surrogate/splits Configuration-performance datasets for TabPFN surrogate training.

Additionally, we host the datasets online, with the corresponding Licenses, source dataset Licenses and corresponding downstream task Licenses:

VLM-Dataset.

LLM-Dataset.

The Croissant RAI metadata files for each dataset are included in this repository:

VLM-Croissant-RAI LLM-Croissant-RAI

Unit testing

To run the unit tests provided in tests/ make sure that pytest is installed.

pip install pytest

Run the tests with the following command:

python -m pytest

Contributing

Contributions are welcome. Please open an issue or submit a pull request.

For usage and licensing terms, see the LICENSE file.

Citations

If you use ScAn-Bench, please cite:

@inproceedings{sermaxhaj2026scanbench,
  title={ScAn-Bench: Evaluating Scaling Analysis Methodology},
  author={A. Sermaxhaj and N. Alipour and D. Sinani and J. Hog and N. Mallik and S. Adriaensen
          and J. Jitsev and D. Stoll},
  booktitle={Advances in Neural Information Processing Systems},
  year={2026},
  note={Evaluations and Datasets Track},
  url={https://arxiv.org/abs/2609.35707}
}

and

@inproceedings{alipour2026tabpfn,
  title={TabPFN-ScAn-Bench: A Surrogate Benchmark for Scaling Analysis Algorithms},
  author={N. Alipour and D. Sinani and A. Sermaxhaj and J. Hog and D. Stoll},
  booktitle={AutoML 2026 non-archival},
  year={2026},
  note={Late-Breaking Abstract}
}

Metadata

Release files for scan-bench 0.2.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scan-bench 0.2.3
File Size Uploaded
scan_bench-0.2.3.tar.gz 23.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for scan-bench 0.2.3
File Interpreter ABI Platform
scan_bench-0.2.3-py3-none-any.whl Python 3 none any Details

Total release size: 47.8 MB

Release files / scan_bench-0.2.3.tar.gz

Download URL scan_bench-0.2.3.tar.gz
Size 23.7 MB
Tags Source
SHA-256 checksum
How to use checksums
2c001d45e5e332b7dba039327dc631e73a7502c9c7fc00989a22e11f88c4910f
BLAKE2b-256 checksum
How to use checksums
22a318d23d7d0a3e866270ed279e3b2c2cedfb3856515ac0fa6046589f520c38
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / scan_bench-0.2.3-py3-none-any.whl

Download URL scan_bench-0.2.3-py3-none-any.whl
Size 24.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
7b22afb7ffc70063d03d41dde5fec9af7a9ed90e79c6a47cf58dac4429450029
BLAKE2b-256 checksum
How to use checksums
39885132bf51f2a0a2415a81c39ee024ee3b2989797d19bfb9bc9e90a647f48e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.2.3 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page