Skip to main content

ScAn-Bench: Evaluating Scaling Analysis Methodology

This repository provides surrogate benchmarks for evaluating scaling analysis (ScAn) methodology on vision-language models (VLMs), large language models (LLMs) and tabular foundation models (TabPFN). The benchmarks approximate the mapping from training configurations to performance, enabling fast evaluation without training full models.

This README describes the ScAn-Bench repository specifically.

For a general overview of the paper, repository structure and artifact maps refer to PaperOverview.

To create a new benchmark based on this framework refer to BenchmarkCreation.

ScAn-Bench Repository

Installation

ScAn-Bench can be used in two ways: as an installed Python package for API querying, or from the cloned source repository for local development and experiment reproduction.

We recommend using a conda environment for both package usage and local development:

conda create -n scan-bench python=3.11
conda activate scan-bench
pip install scan-bench

Install pytorch with CUDA, if you want to utilize the GPU.

Quick start

from scan_bench import TabPFNBenchmark, TabPFNConfig, TabPFNTarget, PerformancePredictorType

bench = TabPFNBenchmark(
    target=TabPFNTarget.VAL_LOSS,                          # or NLL, ROC_AUC
    predictor_type=PerformancePredictorType.ENSEMBLE_XGB,  # or TABPFN, ENSEMBLE_LIGHTGBM, ENSEMBLE_MIX, AUTOGLUON
    device="auto",                                         # or "cpu", "cuda"
)

small = TabPFNConfig(total_cells=2**20, effective_batch_size=16, lr=1e-4,
                     max_features=32, embedding_size=4, num_layers=2, num_datapoints_max=128)
large = TabPFNConfig(total_cells=2**30, effective_batch_size=64, lr=1e-3,
                     max_features=64, embedding_size=64, num_layers=8, num_datapoints_max=256)

print(bench.query(small))                  # {"predictions": {"mean", "uncertainty"}, "model_stats": {"flops", ...}}
print(bench.query_many([small, large]))    # batched, one result per config
print(bench.flops(large), bench.model_params(large))

Package usage

Refer to VLM API, LLM API and TabPFN API for API usages.

Local development and experiment reproduction

For local development and experiment reproduction, use the cloned repository. The training, evaluation, and plotting shell scripts described below are source-tree workflows and should be run from the repository root.

Surrogate training and evaluation

To train and get the performance results for the surrogate benchmarks, run the provided shell scripts. Change DEVICE to 'cuda' in train_surrogates.sh to use the GPU.

VLM pipeline

VLM performance predictor surrogate
bash scan_bench/vlm/performance_surrogate/train/train_surrogates.sh
VLM divergences predictor surrogate
bash scan_bench/vlm/divergence_surrogate/train.sh

LLM pipeline

bash scan_bench/llm/train_surrogates.sh

TabPFN pipeline

bash scan_bench/tabpfn/performance_surrogate/train/train_surrogates.sh

Results

After running the surrogate training scripts, a results directory is created under the corresponding directory for each model family. The generated subdirectories contain JSON files with the evaluation results for the trained surrogate models. These JSON outputs include the metrics reported below and are used to construct the corresponding tables in the paper.

For example:

scan_bench/vlm/performance_surrogate/results/

scan_bench/llm/results/

scan_bench/tabpfn/performance_surrogate/results/

Surrogate Performance (VLM)

Surrogate RMSE ↓ MAE ↓ MDAE ↓ MARPD ↓ R² ↑ R ↑ Corr. ↑
TabPFN 0.21 0.06 0.02 2.69 0.96 0.98 0.98
AutoGluon 0.26 0.12 0.06 5.79 0.94 0.97 0.98
XGB 0.35 0.20 0.12 9.11 0.90 0.95 0.96
Mix 0.37 0.21 0.14 10.11 0.88 0.94 0.95
LGB 0.42 0.29 0.24 14.88 0.86 0.94 0.95

Surrogate Performance (LLM)

Surrogate RMSE ↓ MAE ↓ MDAE ↓ MARPD ↓ R² ↑ R ↑ Corr. ↑
TabPFN 0.27 0.08 0.01 3.02 0.77 0.88 0.95
AutoGluon 0.30 0.11 0.02 4.66 0.72 0.85 0.91
XGB 0.35 0.15 0.04 6.80 0.63 0.80 0.86
Mix 0.33 0.17 0.07 7.65 0.66 0.82 0.86
LGB 0.33 0.17 0.08 7.97 0.66 0.83 0.87

Data

The repository includes pre-collected configuration-performance datasets used to train the surrogate models.

VLM dataset summary

Quantity Count Description
Training configurations 1,535 Total number of collected VLM training configurations.
Failed configurations 134 Configurations that diverged during training.
Successful configurations 1366 Configurations that completed training successfully.
Collected checkpoint rows 8,024 Total checkpoint rows collected across all runs.
Successful checkpoint rows 7,701 Checkpoints from successful runs

For raw logs on the collected VLM data, see the ScAn-VLM-Bench repository.

LLM dataset summary

Quantity Count Description
Training configurations 1,194 Total number of collected LLM training configurations.
Collected checkpoint rows 4,524 Total checkpoints collected across successful and failed runs.
Performance-surrogate rows 4,524 Checkpoints from successful runs used for performance prediction.

TabPFN dataset summary

Quantity Count Description
Training configurations 2,623 Total number of collected TabPFN pretraining configurations (one row per configuration).

Each configuration varies the hyperparameters (lr, effective_batch_size) and the scale parameters (total_cells, embedding_size, num_layers, max_features, num_datapoints_max); see TabPFN search space. Available target is the prior validation loss (val/val_loss).

Data locations

The table below shows where the data is located:

Dataset Path Description
VLM performance data scan_bench/vlm/performance_surrogate/splits Training and test splits for VLM performance surrogate modeling.
VLM divergence data scan_bench/vlm/divergence_surrogate/splits Configuration-level data for predicting failed (diverged) configurations.
LLM performance data scan_bench/llm/splits Configuration-performance datasets for LLM surrogate training.
TabPFN performance data scan_bench/tabpfn/performance_surrogate/splits Configuration-performance datasets for TabPFN surrogate training.

Additionally, we host the datasets online, with the corresponding Licenses, source dataset Licenses and corresponding downstream task Licenses:

VLM-Dataset.

LLM-Dataset.

The Croissant RAI metadata files for each dataset are included in this repository:

VLM-Croissant-RAI LLM-Croissant-RAI

Additional

Unit testing

To run the unit tests provided in tests/ make sure that pytest is installed.

pip install pytest

Run the tests with the following command:

python -m pytest

Contributing

Contributions are welcome. Please open an issue or submit a pull request.

For usage and licensing terms, see the LICENSE file.

Citations

@inproceedings{sermaxhaj2026scanbench,
  title={ScAn-Bench: Evaluating Scaling Analysis Methodology},
  author={A. Sermaxhaj and N. Alipour and D. Sinani and J. Hog and N. Mallik and S. Adriaensen and J. Jitsev and D. Stoll},
  booktitle={Advances in Neural Information Processing Systems},
  year={2026},
  note={Evaluations and Datasets Track},
  url={https://openreview.net/forum?id=EQd9HNVF60}
}

and

@inproceedings{alipour2026tabpfn,
  title={TabPFN-ScAn-Bench: A Surrogate Benchmark for Scaling Analysis Algorithms},
  author={N. Alipour, D. Sinani, A. Sermaxhaj, J. Hog, D. Stoll},
  booktitle={AutoML 2026},
  year={2026},
  note={Late-Breaking Abstract},
}

Metadata

Release files for scan-bench 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scan-bench 0.2.0
File Size Uploaded
scan_bench-0.2.0.tar.gz 23.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for scan-bench 0.2.0
File Interpreter ABI Platform
scan_bench-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 47.8 MB

Release files / scan_bench-0.2.0.tar.gz

Download URL scan_bench-0.2.0.tar.gz
Size 23.7 MB
Tags Source
SHA-256 checksum
How to use checksums
ae7b761afd39e9194c9c34c3cff75ad6e8cd84d4036921b5d95b999660fa978b
BLAKE2b-256 checksum
How to use checksums
07d264126c6b8adb2c37f741f2a6a21808b40702e844cc591fda1f8d30b658c4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release files / scan_bench-0.2.0-py3-none-any.whl

Download URL scan_bench-0.2.0-py3-none-any.whl
Size 24.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
8f3cc2a4e2b1e15a07b4812c9a1ca0d18ae6c715ee482bc79d1f3e8299c80e58
BLAKE2b-256 checksum
How to use checksums
28b4480c72bacb2a02e010a649cacaeccdf0cb0a6e5d0aa0da32f5d62364c6e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release history Release notifications | RSS feed

0.2.3

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page