Skip to main content

SmallMLP

Adaptive neural networks for small nonlinear data, with calibrated conformal intervals.

PyPI version Python License: MIT

SmallMLP is a library for small, nonlinear datasets ($n < 500$). It replaces manual architecture search with adaptive hyperparameters derived from dataset parameters ($n$, $d$, $K$), and produces calibrated prediction intervals through weighted conformal prediction.

  • Regression: 25 / 45 wins, best mean rank (1.82) among 5 models — no tuning.
  • Intervals: 19% narrower than split conformal at equal coverage (31 / 32 wins).
  • Classification: best mean rank among 10 models on 18 datasets; significantly better than LogReg, KNN, SVC; statistically indistinguishable from MLP-128 and RF.

Why SmallMLP?

Standard MLPs need architecture search. Kernel methods need bandwidth tuning. Bayesian methods scale poorly. SmallMLP removes all three problems:

Property Standard MLP GP SmallMLP
Works on small data ($n < 500$) ✗ ✓ ✓
No hyperparameter tuning ✗ ✗ ✓
Prediction intervals / sets ✗ ✓ ✓
Trains in seconds on $n < 500$ ✓ ✗ ✓
Multiclass-aware width ✗ — ✓

Adaptive by construction. Layer width, dropout, and weight decay are computed from dataset parameters — not searched.


Installation

pip install smallmlp

Requires Python 3.10+, torch>=2.0, scikit-learn>=1.3.


Quick start — regression

import numpy as np
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from smallmlp import SmallMLPRegressor

X, y = load_diabetes(return_X_y=True)

# 60 / 20 / 20 split: train / calibration / validation
X_tr, X_tmp, y_tr, y_tmp = train_test_split(X, y, test_size=0.4, random_state=0)
X_cal, X_val, y_cal, y_val = train_test_split(X_tmp, y_tmp, test_size=0.5, random_state=0)

model = SmallMLPRegressor()
model.fit(X_tr, y_tr)
model.fit_conformal(X_cal, y_cal, X_val, y_val, alpha=0.1)

y_hat = model.predict(X_val)
lo, hi = model.predict_interval_conformal(X_val, alpha=0.1)

coverage = np.mean((y_val >= lo) & (y_val <= hi))
print(f"Coverage: {coverage:.3f}  Width: {np.mean(hi - lo):.2f}")

Quick start — classification

import numpy as np
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from smallmlp import SmallMLPClassifier

X, y = load_iris(return_X_y=True)
X_tr, X_tmp, y_tr, y_tmp = train_test_split(X, y, test_size=0.4, stratify=y, random_state=0)
X_cal, X_val, y_cal, y_val = train_test_split(X_tmp, y_tmp, test_size=0.5, stratify=y_tmp, random_state=0)

clf = SmallMLPClassifier(class_weight=None)
clf.fit(X_tr, y_tr)
clf.fit_conformal(X_cal, y_cal, X_val, y_val, alpha=0.1)

probs = clf.predict_proba(X_val)
sets = clf.predict_set(X_val, alpha=0.1)

coverage = np.mean([y_val[i] in sets[i] for i in range(len(y_val))])
avg_size = np.mean([len(s) for s in sets])
print(f"Coverage: {coverage:.3f}  Avg set size: {avg_size:.3f}")

Method

SmallMLP is an adaptive MLP backbone with weighted conformal uncertainty on top.

1. Adaptive width

Layer width is computed from dataset parameters. First layer:

$$ w_1 = \max\left(2K, ; \min\left(\left\lfloor \alpha \cdot \sqrt{n \cdot K \cdot \sqrt{d}} \right\rfloor, ; 4n\right)\right) $$

Subsequent layers:

$$ w_l = \max\left(K, ; \left\lfloor w_1 \cdot \beta^{,l-1} \right\rfloor\right) $$

with $\alpha = 4.0$ and $\beta = 0.7$ chosen by empirical sweep.

Properties:

  • Grows sublinearly with $n$, $K$ (as $\sqrt{\cdot}$), and $d$ (as $d^{1/4}$).
  • First layer bounded below by $2K$; later layers by $K$.
  • Capped at $4n$ to prevent blow-up on tiny datasets.
  • Depth decay $\beta^{l-1}$ shrinks later layers.

For regression, we fall back to a fixed width $\min(2n, 128)$.

2. Adaptive dropout and weight decay

Weight decay is derived from $n$:

$$ \lambda = \frac{1}{n} $$

Dropout uses a clamped function of $d/n$:

$$ p_{\text{drop}} = \min\left(\max\left(3 \cdot \frac{d}{n}, ; 0.1\right), ; 0.5\right) $$

Raw $d/n$ systematically under-estimates required regularization on small data (sweep: optimum in $[0.1, 0.2]$, raw ratio gives $0.02$–$0.08$).

3. Training

Adam optimizer, early stopping on validation loss (20% of train). Class weights are applied only when the imbalance ratio exceeds $2{:}1$.

4. Weighted conformal intervals and sets

The backbone produces an embedding $h(x)$. Given calibration residuals $R_j = |y_j - \hat{y}(x_j)|$ (regression) or scores $s_j = 1 - \hat{p}_{y_j}(x_j)$ (classification), the interval or set for a new $x^*$ uses a weighted quantile of the calibration scores, with weights derived from the kernel:

$$ w_j(x^{\ast}) = \exp\left(-\frac{|h(x^{\ast}) - h(x_j)|^2}{2 h_{\text{cal}}^2}\right) $$

This gives:

  • Finite-sample coverage guarantee under exchangeability.
  • Locally adaptive widths — narrow where data is dense, wide in empty regions.
  • Tuning-free calibration — $h_{\text{cal}}$ selected by grid search on held-out validation.

Results

Regression — point prediction

45 small regression datasets ($n < 500$), 5-fold CV, mean absolute error.

Model Mean rank ↓ Mean MAE ↓ Wins / 45
SmallMLP 1.82 8.43 25
RF (100) 2.33 13.58 14
MLP (256, 128) 3.22 11.57 4
KNN ($k=5$) 3.67 18.89 1
MLP (100,) 3.96 19.07 1

SmallMLP beats MLP (256, 128) by 27% on MAE without any tuning.

Regression — prediction intervals

45 datasets, $\alpha = 0.1$ (target coverage ≥ 0.90), 60/20/20 split.

Method Valid (cov ≥ 0.88) Mean width among valid ↓
Weighted conformal 38 / 45 51.8
Split conformal 32 / 45 64.0
Heuristic $\hat{y} \pm z\sigma$ 17 / 45 145.7

Among 32 datasets where both weighted and split conformal are valid, weighted conformal produces narrower intervals on 31 (mean width ratio 0.81).

Classification — full benchmark

18 small classification datasets, 5-fold CV, 10 models. Metric: mean rank (lower is better, rank 1 = best on that dataset).

Model Mean rank ↓ Mean acc ↑ Wins / 18
SmallMLP (classic) 4.19 0.8120 3
SmallMLP (formula) 4.50 0.8117 2
MLP (128,) 4.64 0.8132 3
MLP (100,) 5.25 0.8101 1
SVC (RBF) 5.31 0.7990 2
MLP (100, 100) 5.56 0.8037 3
RF (100) 5.92 0.8035 6
LogReg 6.17 0.7920 4
KNN ($k=5$) 8.17 0.7715 2

Wilcoxon signed-rank vs SmallMLP (formula), paired across 18 datasets:

Baseline p-value Significant?
MLP (128,) 1.00 no
MLP (100,) 0.62 no
MLP (100,100) 0.26 no
RF (100) 0.37 no
SmallMLP classic 0.62 no
SVC (RBF) 0.047 yes (SmallMLP better)
LogReg 0.013 yes (SmallMLP better)
KNN ($k=5$) 0.002 yes (SmallMLP better)

Summary: SmallMLP achieves the best mean rank among 10 models, is statistically indistinguishable from tuned MLPs and RF, and significantly better than LogReg, KNN, and SVC.

Classification — per-K-bucket accuracy

Bucket SmallMLP (formula) MLP (128,) RF (100)
Binary ($K=2$) 0.8675 0.8697 0.8521
Multiclass ($2<K\le5$) 0.8537 0.8531 0.8210
Hard multiclass ($K>5$) 0.6776 0.6798 0.7019

SmallMLP competitive on binary and moderate multiclass; RF dominates on hard multiclass.

Classification — adaptive dropout ablation

9 small classification datasets, 10-fold CV, accuracy.

Dropout mode Mean accuracy ↑
Unclamped $\min(d/n, ; 0.5)$ 0.8214
Clamped $\min(\max(3d/n, ; 0.1), ; 0.5)$ 0.8269
Fixed $p = 0.2$ 0.8278

Clamped dropout matches fixed $p = 0.2$ within noise; advantage concentrated on multiclass ($K \geq 6$) and high-dimensional ($d \geq 30$) datasets.

Classification — conformal prediction sets

15 small datasets (binary and multiclass), 60/20/20 split, $\alpha = 0.1$.

Dataset $K$ Coverage Avg set size
iris 3 1.00 1.07
wine 3 0.97 1.00
breast cancer 2 0.97 1.00
hepatitis 2 0.90 1.26
parkinsons 2 1.00 1.05
ionosphere 2 0.93 1.06
banknote 2 0.99 1.00
vowel ($K=26$) 26 0.92 2.54

Coverage ≥ 0.90 on the majority of datasets; set size close to 1 on easy problems, grows on hard multiclass ones.


When to use SmallMLP

Good fit:

  • Small datasets ($n < 500$) with nonlinear structure.
  • Scientific instruments: multi-sensor arrays, medical cohorts, chemistry.
  • When uncertainty quantification matters (intervals with coverage guarantee).
  • When you don't want to tune architecture or regularization.

Not a good fit:

  • Large datasets ($n > 10{,}000$) — more data favors standard MLPs.
  • Linear problems — logistic regression or ridge is better.
  • Very high dimensions ($d > 200$) with tiny $n$ — kernel methods suffer.

Use cases

SmallMLP is designed for scientific instruments that produce many features per observation but few observations overall:

  • Telescopes with multi-sensor arrays: 30 sensors × 100–200 observations per campaign.
  • Particle detectors: 10–100 features per event, 200–1000 events per analysis.
  • Medical cohorts: 10–50 clinical variables, 50–300 patients.
  • Chemistry / catalysis: 10–40 descriptors, 50–200 experiments.

In all these settings, standard MLPs overfit, Gaussian Processes scale poorly, and hyperparameter tuning is impractical. SmallMLP provides adaptive regression and classification with calibrated conformal intervals — no tuning required.


Limitations

  • $n < 500$. Fit cost grows with $n$, width, and epochs.
  • Heteroscedastic residuals. On ~15% of regression datasets, weighted conformal undercovers. Attributed to strong heteroscedasticity under non-exchangeability.
  • Linear problems. SmallMLP loses to logistic regression and ridge on linear or near-linear tasks.
  • Raw classification accuracy. SmallMLP matches but does not beat tuned MLPs and RF. Its advantages are tuning-free operation, calibrated uncertainty, and competitive mean rank.
  • Adaptive dropout is only marginally better than fixed $p = 0.2$ on average. Benefit is regime-specific.

Negative results

We document what we tried that did not work, to guide future work:

  • Adaptive depth $L = f(n, d, K)$: no significant improvement over fixed $L = 2$ on $n < 500$.
  • PCA-based effective dimensionality $d_{\text{eff}}$ in the width formula: no improvement over raw $d$.
  • Per-class sample count $\log_2(n/K)$ and softened depth decay $\sqrt{d}/\sqrt{l}$ in the binary width formula: within noise.
  • Bias initialization sweep (zeros, positive, uniform): none beat PyTorch Kaiming default. Correlation between bias negativity and accuracy was spurious.
  • Activation sweep (tanh, GELU, SiLU, algebraic sigmoid, softsign): ReLU wins on 18 small classification datasets; smooth/saturating activations underperform.
  • Sqrt-balanced class weights: improve over linear weights under extreme imbalance, but neither beats no weighting on accuracy.
  • Reweighting / soft-median in conformal scores: does not improve coverage.
  • Heteroscedastic head: degrades point predictions.

What we tried that did work

  • Clamped adaptive dropout min(max(3d/n, 0.1), 0.5) — better than raw d/n.
  • Sqrt-balanced class weights — safer than linear under extreme imbalance (though not better than no weights for accuracy).
  • Weighted conformal — 19% narrower intervals at equal coverage.
  • Multi-class-aware width — competitive with fixed-width baselines without tuning.

API overview

SmallMLPRegressor

SmallMLPRegressor(
    activation="relu",
    lr=1e-3,
    weight_decay=None,   # default 1/n
    max_epochs=500,
    patience=30,
    batch_size=None,     # default min(n, 32)
    val_frac=0.2,
    random_state=42,
    verbose=False,
)

Methods: fit, predict, fit_conformal, predict_interval_conformal.

SmallMLPClassifier

SmallMLPClassifier(
    activation="relu",          # 'relu' | 'tanh' | 'gelu' | 'silu' | 'algsig' | 'softsign'
    lr=1e-3,
    weight_decay=None,          # default 1/n
    max_epochs=500,
    patience=30,
    batch_size=None,
    val_frac=0.2,
    class_weight=None,          # None or "balanced"
    random_state=42,
    verbose=False,
    width_mode="formula",       # "formula" or "classic"
    alpha=4.0,                  # width scale
    beta=0.7,                   # depth decay
    bias_init="kaiming",        # "kaiming" | "zeros" | "positive" | "uniform"
)

Methods: fit, predict, predict_proba, fit_conformal, predict_set, predict_set_labels.


License

MIT License. See LICENSE for details.


Acknowledgments

Built independently during undergraduate studies at Beijing Institute of Technology. Inspired by the author's earlier work on robust gradient boosting for small data (SmallGBM).

Проверь, что в Results таблица согласована с реальными числами — если после обновления кода числа изменились, обнови таблицу.

Metadata

Release files for smallmlp 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for smallmlp 0.3.0
File Size Uploaded
smallmlp-0.3.0.tar.gz 22.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for smallmlp 0.3.0
File Interpreter ABI Platform
smallmlp-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 40.3 kB

Release files / smallmlp-0.3.0.tar.gz

Download URL smallmlp-0.3.0.tar.gz
Size 22.4 kB
Tags Source
SHA-256 checksum
How to use checksums
57aafbe230e85b4e94e627302198f883c2c8c80d6260c4e4225e15e29c7ff768
BLAKE2b-256 checksum
How to use checksums
ec9f50242c28eb8fdbf4ced0f651397b835e88f0bc4f691b37ff2e065e924075
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / smallmlp-0.3.0-py3-none-any.whl

Download URL smallmlp-0.3.0-py3-none-any.whl
Size 17.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f9724630546abe9ee37efe33a01fa304f9a282c42305299005daca0572052bd2
BLAKE2b-256 checksum
How to use checksums
4704264f857c4ba12da19944cbfd1bfb1f35a75d61f04dc0d20465c153876f9c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page