Skip to main content

iLTM: Integrated Large Tabular Model

PyPI License Downloads Python Versions Hugging Face

iLTM is a foundation model for tabular data that integrates tree-derived embeddings, dimensionality-agnostic representations, a meta-trained hypernetwork, multilayer perceptron (MLP) neural networks, and retrieval. iLTM automatically handles feature scaling, categorical features, and missing values.

We release open weights of pre-trained model checkpoints that consistently achieve superior performance across tabular classification and regression tasks, from small to large and high-dimensional tasks.

iLTM architecture diagram

Install

iLTM is accessed through Python. You can install the package via pip:

pip install iltm

iLTM works on Linux, macOS and Windows, and can be executed on CPU and GPU, although GPU is highly recommended for faster execution.

Pre-trained model checkpoints are automatically downloaded from Hugging Face on first use. By default, checkpoints are stored in platform-specific cache directories (e.g., ~/.cache/iltm on Linux, ~/Library/Caches/iltm on macOS). You can specify where model checkpoints are stored by setting the ILTM_CKPT_DIR environment variable:

export ILTM_CKPT_DIR=/path/to/checkpoints

Quick Start

iLTM is designed to be easy to use, with an API similar to scikit-learn.

from iltm import iLTMRegressor, iLTMClassifier

# Regression
reg = iLTMRegressor().fit(X_train, y_train)
y_pred = reg.predict(X_test)

# Classification
clf = iLTMClassifier().fit(X_train, y_train)
proba = clf.predict_proba(X_test)
y_hat = clf.predict(X_test)

# With time limit (returns partial ensemble if time runs out)
reg = iLTMRegressor().fit(X_train, y_train, fit_max_time=3600)  # 1 hour limit

Model Checkpoints

Available checkpoint names:

  • "xgbrconcat" (default): Robust preprocessing + XGBoost embeddings + concatenation
  • "cbrconcat": Robust preprocessing + CatBoost embeddings + concatenation
  • "r128bn": Robust preprocessing with 128-dim bottleneck
  • "rnobn": Robust preprocessing without bottleneck
  • "xgb": XGBoost embeddings only
  • "catb": CatBoost embeddings only
  • "rtr": Robust preprocessing with retrieval
  • "rtrcb": CatBoost embeddings with retrieval

You can also provide a local path to a checkpoint file.

Common key args:

  • checkpoint: checkpoint name or path to model file. Default "xgbrconcat".
  • device: torch device string. Default "cuda:0".
  • n_ensemble: number of generated predictors.
  • batch_size: batch size for weight prediction and inference.
  • preprocessing: "realmlp_td_s_v0" or "minimal" or "none".
  • cat_features: list of categorical column indices.
  • tree_embedding: enable GBDT leaf embeddings.
  • tree_model: "XGBoost_hist" or "CatBoost".
  • concat_tree_with_orig_features: concatenate original features with embeddings.
  • finetuning: end to end finetuning.
  • Retrieval: do_retrieval, retrieval_alpha, retrieval_temperature, retrieval_distance.

Regressor only:

  • clip_predictions: clip to train target range.
  • normalize_predictions: fit a fixed affine output calibration before unscaling.

Classifier only:

  • voting: "soft" or "hard".

Hyperparameter Optimization

iLTM performs best when you tune its hyperparameters.

The package exposes a recommended search space via iltm.get_hyperparameter_search_space, a plain dictionary that maps hyperparameter names to small specs.

The checkpoint parameter is part of this space. It is responsible for selecting one of the built in model checkpoints, which in turn sets other fields such as preprocessing, tree_embedding, and others.

The specification format is intentionally minimal so that it can be re-used in any hyperparameter optimization library or custom tuning procedure.

  • iltm.get_hyperparameter_search_space() gives you the canonical space definition.
  • iltm.sample_hyperparameters(rng) draws a single random configuration from that space for quick baselines and smoke tests.

Development

To run the tests:

pip install -e ".[dev]"
pytest tests/

Citation

If you use iLTM in your research, please cite our paper:

@article{bonet2026iltm,
  title.  = {iLTM: Integrated Large Tabular Model},
  author. = {Bonet, David and Comajoan Cara, Marçal and Calafell, Alvaro and Mas Montserrat, Daniel and Ioannidis, Alexander G.},
  journal = {Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2},
  pages   = {186–197},
  year.   = {2026},
}

License

© Contributors, 2026. Licensed under the Apache-2.0 license.

Metadata

Release files for iltm 0.1.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for iltm 0.1.4
File Size Uploaded
iltm-0.1.4.tar.gz 200.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for iltm 0.1.4
File Interpreter ABI Platform
iltm-0.1.4-py3-none-any.whl Python 3 none any Details

Total release size: 281.3 kB

Release files / iltm-0.1.4.tar.gz

Download URL iltm-0.1.4.tar.gz
Size 200.3 kB
Tags Source
SHA-256 checksum
How to use checksums
5a36654c3e406daf61af11441af72086f2fc4492a42083f932ecf28fb03e214c
BLAKE2b-256 checksum
How to use checksums
74d92b43e50a2b31c6f2de4bd300ae4ed390dffa91b5219e4d8259287db3d87d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.

Transparency log

Release files / iltm-0.1.4-py3-none-any.whl

Download URL iltm-0.1.4-py3-none-any.whl
Size 81.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
15aadc9f7d6457c9bce32c02607925b14b2ea0190079b4e92aea1cd55e051c7b
BLAKE2b-256 checksum
How to use checksums
fd8e4c346a9c3b655098e48a5bad8281341133c00f9dd44e9e265da04861e71a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.4 This release

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page