Skip to main content

DataTunner

Scientific Platform for Optimal Artificial Data Proportion in Deep Learning

DataTunner is a rigorous, reproducible experimentation platform designed to determine the optimal proportion (α*) of artificial data — including augmentation, SMOTE, and CTGAN — in hybrid datasets for neural network training.

Installation

pip install "datatunner[full]"   # everything (torch, SDV, imbalanced-learn)

Mechanism-specific extras keep the base install light:

pip install datatunner            # core only (numpy, pandas, sklearn, matplotlib)
pip install "datatunner[smote]"   # + imbalanced-learn (SMOTE)
pip install "datatunner[ctgan]"   # + SDV / SDMetrics (CTGAN)
pip install "datatunner[image]"   # + torch / torchvision / Pillow (augmentation)

Philosophy

Unlike conventional AutoML tools that treat data proportion as a hyperparameter to be guessed, DataTunner elevates it to a scientific variable with full provenance, statistical fidelity assessment, and isolated experimental environments.

Architecture

datatunner/
├── domain/          # Immutable domain objects
├── generators/      # Abstract generator hierarchy (Augmentation, SMOTE, CTGAN)
├── mixing/          # Isolated mixing engine with leakage detection
├── training/        # Isolated training environments (GPU/seed control)
├── evaluation/      # Quality (fidelity) + Performance (model metrics)
├── optimization/    # Search strategies (Grid, Random, Bayesian)
├── reporting/       # Publication-ready tables and plots
└── infrastructure/  # Seeds, hardware, logging, persistence

Quick Start

import pandas as pd
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
import torch
import torch.nn as nn
from sklearn.metrics import f1_score

from datatunner import DataTunner, ExperimentConfig
from datatunner.domain.generator import GeneratorSpec
from datatunner.generators.smote import SMOTEGenerator
from datatunner.optimization.grid import GridSearch
from datatunner.reporting.exporters import HTMLExporter, LaTeXExporter


# 1. Prepare real data
X, y = make_classification(n_samples=2000, n_features=10, n_classes=2,
                           weights=[0.75, 0.25], random_state=42)
df = pd.DataFrame(X, columns=[f'f{i}' for i in range(10)])
df['target'] = y
train_df, test_df = train_test_split(df, test_size=0.2, random_state=42, stratify=df['target'])


# 2. Define model, train, and evaluate functions
def build_mlp():
    return nn.Sequential(nn.Linear(10, 64), nn.ReLU(),
                         nn.Linear(64, 32), nn.ReLU(), nn.Linear(32, 2))


def train_fn(model, train_data, hyperparams):
    from torch.utils.data import DataLoader, TensorDataset
    X = train_data.drop('target', axis=1).values
    y = train_data['target'].values
    ds = TensorDataset(torch.FloatTensor(X), torch.LongTensor(y))
    loader = DataLoader(ds, batch_size=64, shuffle=True)
    opt = torch.optim.Adam(model.parameters(), lr=1e-3)
    crit = nn.CrossEntropyLoss()
    model.train()
    for _ in range(hyperparams.get('epochs', 20)):
        for bx, by in loader:
            opt.zero_grad()
            loss = crit(model(bx), by)
            loss.backward()
            opt.step()
    return model


def evaluate_fn(model, test_data):
    from datatunner.domain.metrics import ClassificationMetrics
    X = test_data.drop('target', axis=1).values
    y_true = test_data['target'].values
    model.eval()
    with torch.no_grad():
        y_pred = torch.argmax(model(torch.FloatTensor(X)), dim=1).numpy()
    return ClassificationMetrics(f1_macro=f1_score(y_true, y_pred, average='macro'))


# 3. Configure and run experiment
config = ExperimentConfig(
    data_type='tabular',
    search_strategy=GridSearch(),
    alpha_bounds=(0.0, 1.0),
    n_repetitions=3,
    search_budget=6,
    target_metric='f1_macro',
    hyperparams={'epochs': 20, 'batch_size': 64},
)

spec = GeneratorSpec(name="SMOTE_k5", mechanism="smote",
                     hyperparameters={'target_column': 'target', 'k_neighbors': 5},
                     random_state=42)

tunner = DataTunner(config)
report = tunner.run(
    real_data=train_df,
    test_data=test_df,
    model_fn=build_mlp,
    generator=SMOTEGenerator(spec),
    train_fn=train_fn,
    evaluate_fn=evaluate_fn,
)

# 4. Export results
HTMLExporter().export(report, "report.html")
LaTeXExporter().export_tables(report, "./latex_output")

Citation

@software{datatunner2026,
  author = {Rocha, Leandro Costa and Maia de Almeida, Gustavo},
  title = {DataTunner: Optimal Artificial Data Proportion for Deep Learning},
  year = {2026},
  url = {https://github.com/leandrocrx/datatunner}
}

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datatunner-3.2.2.tar.gz (60.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datatunner-3.2.2-py3-none-any.whl (54.2 kB view details)

Uploaded Python 3

File details

Details for the file datatunner-3.2.2.tar.gz.

File metadata

  • Download URL: datatunner-3.2.2.tar.gz
  • Upload date:
  • Size: 60.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.13

File hashes

Hashes for datatunner-3.2.2.tar.gz
Algorithm Hash digest
SHA256 5f6bb0e7c807cbd78f69d222a65fff29efe4f94e534c969757149936a3260cbe
MD5 7076a08fbd596b2f6119f6414f0ae817
BLAKE2b-256 0cbdfce1781ee04ad630414714da74c1140101c894d26e91309b6ba2252336e3

See more details on using hashes here.

File details

Details for the file datatunner-3.2.2-py3-none-any.whl.

File metadata

  • Download URL: datatunner-3.2.2-py3-none-any.whl
  • Upload date:
  • Size: 54.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.13

File hashes

Hashes for datatunner-3.2.2-py3-none-any.whl
Algorithm Hash digest
SHA256 8a6b9effe86c13b38e519044b14e52dba2ca161aaf5265857d60d0934c92b156
MD5 c7e063b89f136ce862350a53659d3313
BLAKE2b-256 6016880eb61709021f30ed8ea5ab44ef1966de8e00ca64eaf84fe065ed1a7aec

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page